Files
tubesync/TODO.md
wimby a48a7cb4f6
All checks were successful
Build TubeSync image / container (push) Successful in 53s
docs: record TubeSync storage and migration state
2026-09-29 00:45:27 +02:00

3.0 KiB

Wimby TubeSync TODO

These tasks track the operational fixes identified while processing the large TubeSync backlog.

Database and queue storage

  • Migrate the TubeSync Django database from SQLite to the existing MariaDB service.

    • MariaDB endpoint: 192.168.1.3:3306 from nc-node1.
    • Current server: MariaDB 11.4.5, utf8mb4.
    • Production database is tubesync with dedicated tubesync@192.168.1.21 credentials.
    • Store the connection string outside Git and set DATABASE_CONNECTION=mysql://... in deployment configuration.
    • Take a verified SQLite backup before migration.
    • Test export/import and Django migrations before switching production.
    • Verify row counts and application health after cutover.
  • Keep all persistent TubeSync runtime state off nc-node1 local ext4.

    • /config, /config/tasks, /config/cache, backups and secrets stay on /data (NFS).
    • /downloads stays on NFS.
    • Deno's transient compiler/module cache uses RAM via DENO_DIR=/dev/shm/deno.
    • No TubeSync queue/cache/migration data is bind-mounted from local ext4.

Backlog cleanup

  • Collapse duplicate download_media_metadata jobs conservatively.

    • Current snapshot: ~19.9k metadata jobs represent only ~2.2k unique media IDs.
    • Keep at most one metadata job per media ID, preferring the highest-priority/latest viable task.
    • Back up the Huey limited-queue DB before modifying it.
    • Validate queue counts and sample task payloads before and after cleanup.
  • Prevent the metadata queue from accumulating duplicate work again.

    • Review remove_duplicates behavior for metadata scheduling.
    • Add a regression test covering repeated scheduling for the same media ID.
    • Prefer deduplication at enqueue time rather than waiting until execution.

YouTube challenge handling

  • Start bgutil-ytdlp-pot-provider and yt-cipher automatically with TubeSync.

    • Both services currently exist in s6 but are not part of the active user bundle.
    • Add the services to the appropriate s6 bundle/dependency graph in the fork.
    • Add a container-level health/smoke check for their local endpoints.
    • Verify yt-dlp no longer reports local provider 502 errors.
  • Improve resilience against YouTube bot/age checks.

    • Decide whether to provide a maintained cookies.txt or another supported authenticated extraction method.
    • Do not commit cookies or credentials to Git.
    • Re-measure metadata failure rate after challenge-provider fixes first.

Throughput

  • Re-measure metadata throughput after database, queue, and provider fixes.

    • Record successful metadata jobs/hour, failures/hour, median task duration, and queue depth.
    • Confirm database is locked failures are eliminated after MariaDB cutover.
  • Consider increasing limited-network metadata concurrency only after stability improves.

    • Start with current worker count as baseline.
    • Increase conservatively and watch YouTube throttling/bot checks, MariaDB load, CPU, and memory.
    • Keep automatic media downloads below metadata priority.