From 91fbd63a313fcdd9266ff7fd7a9b4697ae5210a6 Mon Sep 17 00:00:00 2001 From: wimby Date: Sun, 27 Sep 2026 18:00:35 +0200 Subject: [PATCH] docs: track TubeSync operational follow-ups --- TODO.md | 57 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++ WIMBY.md | 2 ++ 2 files changed, 59 insertions(+) create mode 100644 TODO.md diff --git a/TODO.md b/TODO.md new file mode 100644 index 00000000..202acf84 --- /dev/null +++ b/TODO.md @@ -0,0 +1,57 @@ +# Wimby TubeSync TODO + +These tasks track the operational fixes identified while processing the large TubeSync backlog. + +## Database and queue storage + +- [ ] Migrate the TubeSync Django database from SQLite to the existing MariaDB service. + - MariaDB endpoint: `192.168.1.3:3306` from `nc-node1`. + - Current server: MariaDB 11.4.5, `utf8mb4`. + - Create a dedicated `tubesync` database and least-privilege user. + - Store the connection string outside Git and set `DATABASE_CONNECTION=mysql://...` in deployment configuration. + - Take a verified SQLite backup before migration. + - Test export/import and Django migrations before switching production. + - Verify row counts and application health after cutover. + +- [ ] Move Huey SQLite queue files off NFS onto local ext4 storage on `nc-node1`. + - Keep `/downloads` on NFS. + - Bind-mount a local directory onto `/config/tasks`. + - Stop TubeSync before copying the queue DB/WAL/SHM files. + - Verify all four Huey consumers restart and queue counts are preserved. + +## Backlog cleanup + +- [ ] Collapse duplicate `download_media_metadata` jobs conservatively. + - Current snapshot: ~19.9k metadata jobs represent only ~2.2k unique media IDs. + - Keep at most one metadata job per media ID, preferring the highest-priority/latest viable task. + - Back up the Huey limited-queue DB before modifying it. + - Validate queue counts and sample task payloads before and after cleanup. + +- [ ] Prevent the metadata queue from accumulating duplicate work again. + - Review `remove_duplicates` behavior for metadata scheduling. + - Add a regression test covering repeated scheduling for the same media ID. + - Prefer deduplication at enqueue time rather than waiting until execution. + +## YouTube challenge handling + +- [ ] Start `bgutil-ytdlp-pot-provider` and `yt-cipher` automatically with TubeSync. + - Both services currently exist in s6 but are not part of the active `user` bundle. + - Add the services to the appropriate s6 bundle/dependency graph in the fork. + - Add a container-level health/smoke check for their local endpoints. + - Verify yt-dlp no longer reports local provider 502 errors. + +- [ ] Improve resilience against YouTube bot/age checks. + - Decide whether to provide a maintained `cookies.txt` or another supported authenticated extraction method. + - Do not commit cookies or credentials to Git. + - Re-measure metadata failure rate after challenge-provider fixes first. + +## Throughput + +- [ ] Re-measure metadata throughput after database, queue, and provider fixes. + - Record successful metadata jobs/hour, failures/hour, median task duration, and queue depth. + - Confirm `database is locked` failures are eliminated after MariaDB cutover. + +- [ ] Consider increasing limited-network metadata concurrency only after stability improves. + - Start with current worker count as baseline. + - Increase conservatively and watch YouTube throttling/bot checks, MariaDB load, CPU, and memory. + - Keep automatic media downloads below metadata priority. diff --git a/WIMBY.md b/WIMBY.md index 0fac99b1..b5d84389 100644 --- a/WIMBY.md +++ b/WIMBY.md @@ -33,3 +33,5 @@ A push to `main` builds and publishes: The deployment should normally track `:latest`; the SHA tag is available for rollback. The Gitea Actions runner builds this image natively on ARM64. + +Operational follow-up work is tracked in [`TODO.md`](TODO.md).