docs: track TubeSync operational follow-ups
All checks were successful
Build TubeSync image / container (push) Successful in 53s

This commit is contained in:
wimby
2026-09-27 18:00:35 +02:00
parent 1c821ba1f7
commit 91fbd63a31
2 changed files with 59 additions and 0 deletions

57
TODO.md Normal file
View File

@@ -0,0 +1,57 @@
# Wimby TubeSync TODO
These tasks track the operational fixes identified while processing the large TubeSync backlog.
## Database and queue storage
- [ ] Migrate the TubeSync Django database from SQLite to the existing MariaDB service.
- MariaDB endpoint: `192.168.1.3:3306` from `nc-node1`.
- Current server: MariaDB 11.4.5, `utf8mb4`.
- Create a dedicated `tubesync` database and least-privilege user.
- Store the connection string outside Git and set `DATABASE_CONNECTION=mysql://...` in deployment configuration.
- Take a verified SQLite backup before migration.
- Test export/import and Django migrations before switching production.
- Verify row counts and application health after cutover.
- [ ] Move Huey SQLite queue files off NFS onto local ext4 storage on `nc-node1`.
- Keep `/downloads` on NFS.
- Bind-mount a local directory onto `/config/tasks`.
- Stop TubeSync before copying the queue DB/WAL/SHM files.
- Verify all four Huey consumers restart and queue counts are preserved.
## Backlog cleanup
- [ ] Collapse duplicate `download_media_metadata` jobs conservatively.
- Current snapshot: ~19.9k metadata jobs represent only ~2.2k unique media IDs.
- Keep at most one metadata job per media ID, preferring the highest-priority/latest viable task.
- Back up the Huey limited-queue DB before modifying it.
- Validate queue counts and sample task payloads before and after cleanup.
- [ ] Prevent the metadata queue from accumulating duplicate work again.
- Review `remove_duplicates` behavior for metadata scheduling.
- Add a regression test covering repeated scheduling for the same media ID.
- Prefer deduplication at enqueue time rather than waiting until execution.
## YouTube challenge handling
- [ ] Start `bgutil-ytdlp-pot-provider` and `yt-cipher` automatically with TubeSync.
- Both services currently exist in s6 but are not part of the active `user` bundle.
- Add the services to the appropriate s6 bundle/dependency graph in the fork.
- Add a container-level health/smoke check for their local endpoints.
- Verify yt-dlp no longer reports local provider 502 errors.
- [ ] Improve resilience against YouTube bot/age checks.
- Decide whether to provide a maintained `cookies.txt` or another supported authenticated extraction method.
- Do not commit cookies or credentials to Git.
- Re-measure metadata failure rate after challenge-provider fixes first.
## Throughput
- [ ] Re-measure metadata throughput after database, queue, and provider fixes.
- Record successful metadata jobs/hour, failures/hour, median task duration, and queue depth.
- Confirm `database is locked` failures are eliminated after MariaDB cutover.
- [ ] Consider increasing limited-network metadata concurrency only after stability improves.
- Start with current worker count as baseline.
- Increase conservatively and watch YouTube throttling/bot checks, MariaDB load, CPU, and memory.
- Keep automatic media downloads below metadata priority.

View File

@@ -33,3 +33,5 @@ A push to `main` builds and publishes:
The deployment should normally track `:latest`; the SHA tag is available for rollback.
The Gitea Actions runner builds this image natively on ARM64.
Operational follow-up work is tracked in [`TODO.md`](TODO.md).