Skip to main content
The repository contains the Next.js web application, Cloudflare Worker platform, YouTube extraction package, processor container, and this documentation project.
The combined local command starts the platform and web application with local Cloudflare storage. Local dashboard requests use isolated demo identities; they do not create resources in production D1. To call private platform endpoints without an API key or demo header, start the platform with the explicit local authentication bypass:
The platform development scripts use src/index.local.ts. This entry point recreates a warm YouTube processor container when the Worker reloads, because reused containers can lose outbound connectivity after a Wrangler reload. Local database and cache storage are preserved. An in-flight provider request may need to be retried after saving code. Production uses src/index.ts and keeps its normal container lifecycle. The bypass is disabled by default. It activates only for loopback requests when ENVIRONMENT is not production and uses one stable local demo account. Run the documentation separately with:
Workers AI and explicitly remote bindings can incur usage even during local development. Never commit .dev.vars, API keys, OAuth secrets, or email credentials.

Google and GitHub sign-in

The login page offers Google and GitHub for both registration and returning users. Set GOOGLE_CLIENT_ID, GOOGLE_CLIENT_SECRET, GITHUB_CLIENT_ID, and GITHUB_CLIENT_SECRET in the untracked platform/.dev.vars file. These credentials belong to the platform, never the browser bundle. Create a GitHub OAuth App with homepage http://localhost:3000 and authorization callback http://localhost:3000/api/auth/callback/github. Register http://localhost:3000/api/auth/callback/google with your Google OAuth web client. Better Auth requests GitHub’s read:user and user:email scopes, including access to private email addresses for sign-in. For production, use separate provider credentials and replace the local origin in each callback URL with the configured AUTH_BASE_URL. Store the credentials as Worker secrets. The login page preserves the requested dashboard destination after authentication. The existing magic-link endpoint remains available for the CLI device authorization page and previously issued email links. It is no longer offered on /login.

Agent smoke test

Start the local Worker with Docker running and local database migrations applied. To use the local demo account for this test:
These command-line overrides apply to the local process. Storage stays local, while the remote Workers AI binding runs GLM. Do not add --local, which disables that binding. Production remains configured for the Agent tester allowlist. In another terminal:
The smoke test explicitly requests the legacy format to inspect detailed evidence, and calls the real Worker API for research and video inspection, using the shared GLM configuration for the agent and its analysts. For the default recommendation query, it requires comparative routing and completed transcript analyses for four distinct videos. It requires full answers without PARTIAL_EVIDENCE, citations, evidence charges, and storyboard citations and analysis in the inspect result. It polls every three seconds and backs off on rate limits without creating another run. Each run reserves 22 available credits and refunds unused credits after completion, failure, or cancellation. Successful evidence operations remain chargeable if the run fails. For authenticated testing, provide AGENT_TEST_TOKEN. To test the Agent allowlist, apply local D1 migrations, run the local Worker in AGENT_ACCESS_MODE:allowlist, add the verified tester email to agent_access_allowlist, and also provide AGENT_TEST_UNLISTED_TOKEN for an account absent from the allowlist. Agent access is separate from ADMIN_EMAILS_SECRET. The test reports the denial check as skipped when that second token is absent. Keep tokens in environment variables, outside source control.
Classification selects the route, storyboard access, and researchVideoCount before research begins. New research decisions choose 1 to 8 videos; inspection always selects 1. researchBreadth describes the task rather than fixing the count. A separate requiredVideoCount records only an explicit user request for a source count, not an answer-item count. Historical decisions without a count retain their former two/four-video defaults. Research submits selected transcripts together with at most four active analysts; inspection retains two active analysts. Each video has independent persistence and billing. Classification has 20 seconds, ordinary research has 40 seconds, visual research has 120 seconds, and finalization has 60 seconds. Phase deadlines are persisted across recovery. The visual window accommodates frame transport of up to 70 seconds plus up to 20 seconds of image analysis. The transcript analysis budget follows the classified count. Larger counts are not a completion guarantee within the research deadline. Completed compact results expose coverage.targetVideos and coverage.reviewedVideos, plus coverage.requiredVideos when explicitly requested. These counts measure distinct videos with usable transcript evidence. Missing an internal target alone does not produce partial. Missing a user-required source count adds PARTIAL_EVIDENCE. Source caveats remain separate from answer completeness. Transcript analysts return at most five relevant findings per video. This bounds extraction work without fixing the number of recommendations in the final answer. When the agent selects forced finalization, it exits the research phase before starting synthesis. Synthesis receives one continuous budget of up to 60 seconds. Context gathering gets at most 10 seconds. The first answer reserves up to 20 seconds for one repair; a resumed phase with less time divides its remaining time between the two attempts. Saving the validated result and settling billing run outside the model-processing windows, with a separate 30-second persistence timeout. Billing settlement is idempotent and retried by the durable watchdog if needed. Total completion time includes queueing, up to 120 seconds of ordinary processing or 200 seconds when visual research is enabled, and persistence. These budgets do not guarantee provider success. Crossing the research cutoff does not cancel and restart that synthesis call. Classification uses a required structured tool call to choose scope, research breadth, and visual evidence access. Requests outside YouTube video research and synthesis, including tasks requiring unsupported video platforms, produce route: rejected with a reason. These complete with compact outcome: rejected or legacy intent: rejected and an OUT_OF_SCOPE warning, without fetching evidence or running downstream models. Ambiguous supported requests can still request clarification. New executable routes must include useStoryboard. Visual questions about slides, charts, interfaces, scenes, or demonstrations enable it; ordinary transcript summaries and verbal comparisons normally disable it. When false, the research model receives neither get_video_storyboard nor get_video_frames, and access to both visual providers is disabled. The choice is persisted for recovery. Legacy routes without the field retain their existing tool access. Answer blocks accept up to 12 validated evidence references to support comparisons across four videos. The classifier also supplies a searchQuery, which the application executes before the first research model step. Recovery retains the consumed search budget. This removes a separate model call to plan the same search. Classifier and answer tool schemas expose top-level object fields, with conditional routing and citation validation enforced in application code. If transcript requests fail and only discovery metadata remains, the agent skips unsupported recommendation synthesis and returns an explicit NO_CONTENT_EVIDENCE warning with source links. It does not quote promotional descriptions as an answer. Recommendation synthesis prioritizes a short list of practical tasks, examples, and source-attributed claims; successful citations establish traceability, not independent verification of a video’s claims. To repeat a specific research query locally alongside the inspect regression:
For custom queries, the smoke test checks completion and citations; review the answer and result.coverage separately. The default recommendation query explicitly requests four source videos and checks that requirement.

Compact agent responses

Compact responses are the default for admission and polling. Existing clients that parse the previous detailed contract must add ?responseFormat=legacy to both URLs. The format is a per-request presentation option; selecting legacy at admission does not change later polling responses. Both formats read the same stored run; changing formats does not rerun or charge for analysis. In Postman:
  1. Send POST {{baseUrl}}/v1/agent?responseFormat=compact with your usual authorization and JSON body such as {"message":"Research the best design skills for frontend developers using Claude Code"}. Use the returned session and run IDs to retrieve existing work. Each accepted POST creates a new run.
  2. Save runId and sessionId from the 202 receipt.
  3. Poll GET {{baseUrl}}/v1/agent/{{sessionId}}/runs/{{runId}}?responseFormat=compact until the status is terminal.
  4. Read result.answer and resolve its [1] references using result.sources. Billing appears once at the top level. Keep agentMessageId if you want to use that answer as a follow-up’s parentMessageId.
Research and inspect use the same compact schema. Each cited video appears once in sources, numbered by its first reference in the answer. Channel and playlist evidence retains its source identity when a video is not involved. Source URLs come from persisted evidence or the known video ID. Several adjacent excerpts from one video collapse to one source reference. Excerpts and timestamps remain stored internally. status describes execution. A completed run has one of these result.outcome values: Both compact and legacy run responses include request.message, copied from the persisted input on admission and polling. It is not regenerated by a model. Warnings remain in result.warnings. New transcript warnings include videoId so source-specific caveats do not appear to apply to every video. Provider errors and other caveats remain visible even when they do not change the outcome. Pending, running, failed, and cancelled responses have no result or billing unless a persisted result exists; failed runs retain their error message. Legacy RESEARCH_COVERAGE_SHORTFALL warnings no longer force a compact result to partial and are omitted from compact warnings; their coverage is derived from stored route and analysis metadata when available. Other historical runs use their persisted warnings, so older results without NO_CONTENT_EVIDENCE cannot always distinguish missing content from other partial results. For additional details, append &include=evidence,artifacts,diagnostics. Request only the groups you need:
  • evidence adds result.evidence with original excerpt IDs, text, and available timestamps. Each entry’s sourceId points to result.sources[].id.
  • artifacts adds result.artifacts with stored tool output artifacts.
  • diagnostics adds top-level routing, user message ID, conversation turn, model step count, tool call count, and transcriptAnalysis attempt records. Classifier and isolated analyst calls are excluded from the existing model step counter. Transcript records identify validation rules, zero-based finding and field indexes, exact repair feedback, finish reason, token usage, timing, and cancellation reason. Rejected model text and the referenced source windows are private diagnostic evidence, never accepted answer evidence. These captures are bounded to 24,000 output characters, 15 source windows, 100 issues, and 4,000 repair-feedback characters; captureTruncated identifies a shortened capture. Ordinary logs contain only rule codes and execution metadata. Legacy run responses expose the records as transcriptDiagnostics. Older runs without these records cannot recover the original rejection reason retroactively.
Invalid formats or include values return 422. include works with the default compact format and is rejected with responseFormat=legacy. Session restoration retains its existing format. The compact response is a serialization change; transcript analysis and exact internal citation validation remain in place.

Adaptive answer scope

Both research and inspect honor the requested number of points and level of detail. Without an explicit count, the model chooses a useful number from the evidence. Recommendation count is separate from the two- or four-video research target. The checked-in configuration uses GLM on Fireworks for classification and research, with DeepSeek V4 Flash 0731 on Fireworks for finalization. Set FIREWORKS_API_KEY in the untracked .dev.vars. Local overrides should use AGENT_FINALIZER_PROVIDER=fireworks and AGENT_FINALIZER_MODEL=deepseek-v4-flash-0731. The native adapter also supports deepseek-v4-flash-0731 with bounded thinking and gpt-oss-120b with low reasoning effort. It adds 1,024 tokens of reasoning headroom to the answer allowance; Fireworks counts both reasoning and answer tokens inside the combined ceiling. GPT-OSS low reasoning is not a fixed reasoning-token budget. AGENT_GLM_PROVIDER=fireworks is the default for GLM classification, research, transcript extraction, visual analysis and research tool-argument repair. Set AGENT_GLM_PROVIDER=workers-ai to switch those roles back to the retained Workers AI binding and AI Gateway. Finalizer selection is independent. Fireworks GLM uses native low reasoning for these roles with the existing research token ceilings, and the SDK preserves reasoning between tool calls. Research completion hands off to the configured finalizer under its separate deadline. A terminal tool request or malformed terminal arguments trigger that handoff; research drafts are not persisted as the final answer. Keep test captures under ignored .scratch/; do not commit credentials, prompts, raw responses or benchmark results. Normal and recovery synthesis share concise writing guidance and native output limits. The classifier selects answerDetail: standard allows 1,500 output tokens per generation and detailed allows 2,500. Legacy routes default to standard. These SDK ceilings apply to research generations that can answer and terminal-tool argument repair. The unified structured finalizer allows 3,000 tokens for standard answers or 4,000 for detailed answers, plus 1,000 additional tokens on its one repair attempt. The Fireworks adapter adds its separate 1,024-token reasoning allowance. Answer blocks allow up to 20 sections, with the existing 20,000-character validation limit. Large lists may group related items into sections. Transcript analysts return up to five concise findings with exact evidence references, aiming for 250 to 350 output tokens with a 1,200-token ceiling. They read the full transcript but do not generate a separate summary. Each claim retains relevant attribution and caveats, with up to three short source warnings. The persisted summary field is an application-generated finding count for compatibility. Recovery uses the selected output-token ceiling within the shared 60-second finalization phase; a validation repair does not restart that clock. The local smoke test allows 240 seconds for queueing, visual processing, persistence, and polling. These limits bound resource usage rather than guarantee any requested count. If evidence or the response budget prevents fulfilling the request, the model is instructed to explain the shortfall using ANSWER_SCOPE_SHORTFALL, which the application maps to PARTIAL_EVIDENCE. Source caveats use SOURCE_CAVEAT and do not by themselves make the outcome partial. Runtime timeout warnings and coverage counts remain application-owned. Recommendations should have distinct categories, practical examples, and contextualized quantitative claims. Not analyzing every search result is not itself a coverage gap; actual coverage shortfalls and material source limitations remain visible.

Shared video catalog

The platform also uses local VIDEO_CATALOG D1 and VIDEO_ASSETS R2 bindings. npm --prefix platform run db:migrate:local applies both the account and video catalog migrations. The catalog stores requested public video assets incrementally; see reference/engineering/VIDEO_CATALOG.md for the storage layout and deployment requirements. The checked-in catalog IDs refer to provisioned hosted databases; --local and the catalog binding’s remote: false keep local development on local storage. Self-hosted deployments must configure their own databases and buckets.

Worker extraction

Core YouTube extraction runs in the Worker by default with YOUTUBE_EXTRACTION_BACKEND=worker. Install shared library dependencies with npm ci --prefix packages/all-things-youtube before starting Wrangler. Configure OUTBOUND_PROXY_URLS or OUTBOUND_PROXY_URL in platform/.dev.vars for proxy fallback. Storyboards and exact frames still require Docker. Set YOUTUBE_EXTRACTION_BACKEND=container to use the rollback backend. See the repository document reference/engineering/WORKER_EXTRACTION.md for retry limits and rollout.