src/index.local.ts. This entry point recreates a warm YouTube processor container when the Worker reloads, because reused containers can lose outbound connectivity after a Wrangler reload. Local database and cache storage are preserved. An in-flight provider request may need to be retried after saving code. Production uses src/index.ts and keeps its normal container lifecycle.
The bypass is disabled by default. It activates only for loopback requests when
ENVIRONMENT is not production and uses one stable local demo account.
Run the documentation separately with:
Google and GitHub sign-in
The login page offers Google and GitHub for both registration and returning users. SetGOOGLE_CLIENT_ID, GOOGLE_CLIENT_SECRET, GITHUB_CLIENT_ID, and GITHUB_CLIENT_SECRET in the untracked platform/.dev.vars file. These credentials belong to the platform, never the browser bundle.
Create a GitHub OAuth App with homepage http://localhost:3000 and authorization callback http://localhost:3000/api/auth/callback/github. Register http://localhost:3000/api/auth/callback/google with your Google OAuth web client. Better Auth requests GitHub’s read:user and user:email scopes, including access to private email addresses for sign-in.
For production, use separate provider credentials and replace the local origin in each callback URL with the configured AUTH_BASE_URL. Store the credentials as Worker secrets. The login page preserves the requested dashboard destination after authentication.
The existing magic-link endpoint remains available for the CLI device authorization page and previously issued email links. It is no longer offered on /login.
Agent smoke test
Start the local Worker with Docker running and local database migrations applied. To use the local demo account for this test:--local, which disables that binding. Production remains configured for the Agent tester allowlist.
In another terminal:
PARTIAL_EVIDENCE, citations, evidence charges, and storyboard citations and analysis in the inspect result. It polls every three seconds and backs off on rate limits without creating another run. Each run reserves 22 available credits and refunds unused credits after completion, failure, or cancellation. Successful evidence operations remain chargeable if the run fails.
For authenticated testing, provide AGENT_TEST_TOKEN. To test the Agent allowlist, apply local D1 migrations, run the local Worker in AGENT_ACCESS_MODE:allowlist, add the verified tester email to agent_access_allowlist, and also provide AGENT_TEST_UNLISTED_TOKEN for an account absent from the allowlist. Agent access is separate from ADMIN_EMAILS_SECRET. The test reports the denial check as skipped when that second token is absent. Keep tokens in environment variables, outside source control.
researchVideoCount before research begins. New research decisions choose 1 to 8 videos; inspection always selects 1. researchBreadth describes the task rather than fixing the count. A separate requiredVideoCount records only an explicit user request for a source count, not an answer-item count. Historical decisions without a count retain their former two/four-video defaults. Research submits selected transcripts together with at most four active analysts; inspection retains two active analysts. Each video has independent persistence and billing. Classification has 20 seconds, ordinary research has 40 seconds, visual research has 120 seconds, and finalization has 60 seconds. Phase deadlines are persisted across recovery. The visual window accommodates frame transport of up to 70 seconds plus up to 20 seconds of image analysis. The transcript analysis budget follows the classified count. Larger counts are not a completion guarantee within the research deadline.
Completed compact results expose coverage.targetVideos and coverage.reviewedVideos, plus coverage.requiredVideos when explicitly requested. These counts measure distinct videos with usable transcript evidence. Missing an internal target alone does not produce partial. Missing a user-required source count adds PARTIAL_EVIDENCE. Source caveats remain separate from answer completeness.
Transcript analysts return at most five relevant findings per video. This bounds extraction work without fixing the number of recommendations in the final answer. When the agent selects forced finalization, it exits the research phase before starting synthesis. Synthesis receives one continuous budget of up to 60 seconds. Context gathering gets at most 10 seconds. The first answer reserves up to 20 seconds for one repair; a resumed phase with less time divides its remaining time between the two attempts. Saving the validated result and settling billing run outside the model-processing windows, with a separate 30-second persistence timeout. Billing settlement is idempotent and retried by the durable watchdog if needed. Total completion time includes queueing, up to 120 seconds of ordinary processing or 200 seconds when visual research is enabled, and persistence. These budgets do not guarantee provider success. Crossing the research cutoff does not cancel and restart that synthesis call.
Classification uses a required structured tool call to choose scope, research breadth, and visual evidence access. Requests outside YouTube video research and synthesis, including tasks requiring unsupported video platforms, produce route: rejected with a reason. These complete with compact outcome: rejected or legacy intent: rejected and an OUT_OF_SCOPE warning, without fetching evidence or running downstream models. Ambiguous supported requests can still request clarification.
New executable routes must include useStoryboard. Visual questions about slides, charts, interfaces, scenes, or demonstrations enable it; ordinary transcript summaries and verbal comparisons normally disable it. When false, the research model receives neither get_video_storyboard nor get_video_frames, and access to both visual providers is disabled. The choice is persisted for recovery. Legacy routes without the field retain their existing tool access. Answer blocks accept up to 12 validated evidence references to support comparisons across four videos.
The classifier also supplies a searchQuery, which the application executes before the first research model step. Recovery retains the consumed search budget. This removes a separate model call to plan the same search. Classifier and answer tool schemas expose top-level object fields, with conditional routing and citation validation enforced in application code.
If transcript requests fail and only discovery metadata remains, the agent skips unsupported recommendation synthesis and returns an explicit NO_CONTENT_EVIDENCE warning with source links. It does not quote promotional descriptions as an answer. Recommendation synthesis prioritizes a short list of practical tasks, examples, and source-attributed claims; successful citations establish traceability, not independent verification of a video’s claims.
To repeat a specific research query locally alongside the inspect regression:
result.coverage separately. The default recommendation query explicitly requests four source videos and checks that requirement.
Compact agent responses
Compact responses are the default for admission and polling. Existing clients that parse the previous detailed contract must add?responseFormat=legacy to both URLs. The format is a per-request presentation option; selecting legacy at admission does not change later polling responses. Both formats read the same stored run; changing formats does not rerun or charge for analysis.
In Postman:
- Send
POST {{baseUrl}}/v1/agent?responseFormat=compactwith your usual authorization and JSON body such as{"message":"Research the best design skills for frontend developers using Claude Code"}. Use the returned session and run IDs to retrieve existing work. Each accepted POST creates a new run. - Save
runIdandsessionIdfrom the202receipt. - Poll
GET {{baseUrl}}/v1/agent/{{sessionId}}/runs/{{runId}}?responseFormat=compactuntil the status is terminal. - Read
result.answerand resolve its[1]references usingresult.sources. Billing appears once at the top level. KeepagentMessageIdif you want to use that answer as a follow-up’sparentMessageId.
sources, numbered by its first reference in the answer. Channel and playlist evidence retains its source identity when a video is not involved. Source URLs come from persisted evidence or the known video ID. Several adjacent excerpts from one video collapse to one source reference. Excerpts and timestamps remain stored internally.
status describes execution. A completed run has one of these result.outcome values:
Both compact and legacy run responses include
request.message, copied from the persisted input on admission and polling. It is not regenerated by a model.
Warnings remain in result.warnings. New transcript warnings include videoId so source-specific caveats do not appear to apply to every video. Provider errors and other caveats remain visible even when they do not change the outcome. Pending, running, failed, and cancelled responses have no result or billing unless a persisted result exists; failed runs retain their error message. Legacy RESEARCH_COVERAGE_SHORTFALL warnings no longer force a compact result to partial and are omitted from compact warnings; their coverage is derived from stored route and analysis metadata when available. Other historical runs use their persisted warnings, so older results without NO_CONTENT_EVIDENCE cannot always distinguish missing content from other partial results.
For additional details, append &include=evidence,artifacts,diagnostics. Request only the groups you need:
evidenceaddsresult.evidencewith original excerpt IDs, text, and available timestamps. Each entry’ssourceIdpoints toresult.sources[].id.artifactsaddsresult.artifactswith stored tool output artifacts.diagnosticsadds top-level routing, user message ID, conversation turn, model step count, tool call count, andtranscriptAnalysisattempt records. Classifier and isolated analyst calls are excluded from the existing model step counter. Transcript records identify validation rules, zero-based finding and field indexes, exact repair feedback, finish reason, token usage, timing, and cancellation reason. Rejected model text and the referenced source windows are private diagnostic evidence, never accepted answer evidence. These captures are bounded to 24,000 output characters, 15 source windows, 100 issues, and 4,000 repair-feedback characters;captureTruncatedidentifies a shortened capture. Ordinary logs contain only rule codes and execution metadata. Legacy run responses expose the records astranscriptDiagnostics. Older runs without these records cannot recover the original rejection reason retroactively.
422. include works with the default compact format and is rejected with responseFormat=legacy. Session restoration retains its existing format. The compact response is a serialization change; transcript analysis and exact internal citation validation remain in place.
Adaptive answer scope
Both research and inspect honor the requested number of points and level of detail. Without an explicit count, the model chooses a useful number from the evidence. Recommendation count is separate from the two- or four-video research target. The checked-in configuration uses GLM on Fireworks for classification and research, with DeepSeek V4 Flash 0731 on Fireworks for finalization. SetFIREWORKS_API_KEY in the untracked .dev.vars. Local overrides should use AGENT_FINALIZER_PROVIDER=fireworks and AGENT_FINALIZER_MODEL=deepseek-v4-flash-0731. The native adapter also supports deepseek-v4-flash-0731 with bounded thinking and gpt-oss-120b with low reasoning effort. It adds 1,024 tokens of reasoning headroom to the answer allowance; Fireworks counts both reasoning and answer tokens inside the combined ceiling. GPT-OSS low reasoning is not a fixed reasoning-token budget. AGENT_GLM_PROVIDER=fireworks is the default for GLM classification, research, transcript extraction, visual analysis and research tool-argument repair. Set AGENT_GLM_PROVIDER=workers-ai to switch those roles back to the retained Workers AI binding and AI Gateway. Finalizer selection is independent. Fireworks GLM uses native low reasoning for these roles with the existing research token ceilings, and the SDK preserves reasoning between tool calls. Research completion hands off to the configured finalizer under its separate deadline. A terminal tool request or malformed terminal arguments trigger that handoff; research drafts are not persisted as the final answer. Keep test captures under ignored .scratch/; do not commit credentials, prompts, raw responses or benchmark results.
Normal and recovery synthesis share concise writing guidance and native output limits. The classifier selects answerDetail: standard allows 1,500 output tokens per generation and detailed allows 2,500. Legacy routes default to standard. These SDK ceilings apply to research generations that can answer and terminal-tool argument repair. The unified structured finalizer allows 3,000 tokens for standard answers or 4,000 for detailed answers, plus 1,000 additional tokens on its one repair attempt. The Fireworks adapter adds its separate 1,024-token reasoning allowance. Answer blocks allow up to 20 sections, with the existing 20,000-character validation limit. Large lists may group related items into sections. Transcript analysts return up to five concise findings with exact evidence references, aiming for 250 to 350 output tokens with a 1,200-token ceiling. They read the full transcript but do not generate a separate summary. Each claim retains relevant attribution and caveats, with up to three short source warnings. The persisted summary field is an application-generated finding count for compatibility. Recovery uses the selected output-token ceiling within the shared 60-second finalization phase; a validation repair does not restart that clock. The local smoke test allows 240 seconds for queueing, visual processing, persistence, and polling.
These limits bound resource usage rather than guarantee any requested count. If evidence or the response budget prevents fulfilling the request, the model is instructed to explain the shortfall using ANSWER_SCOPE_SHORTFALL, which the application maps to PARTIAL_EVIDENCE. Source caveats use SOURCE_CAVEAT and do not by themselves make the outcome partial. Runtime timeout warnings and coverage counts remain application-owned. Recommendations should have distinct categories, practical examples, and contextualized quantitative claims. Not analyzing every search result is not itself a coverage gap; actual coverage shortfalls and material source limitations remain visible.
Shared video catalog
The platform also uses localVIDEO_CATALOG D1 and VIDEO_ASSETS R2 bindings. npm --prefix platform run db:migrate:local applies both the account and video catalog migrations. The catalog stores requested public video assets incrementally; see reference/engineering/VIDEO_CATALOG.md for the storage layout and deployment requirements. The checked-in catalog IDs refer to provisioned hosted databases; --local and the catalog binding’s remote: false keep local development on local storage. Self-hosted deployments must configure their own databases and buckets.
Worker extraction
Core YouTube extraction runs in the Worker by default withYOUTUBE_EXTRACTION_BACKEND=worker. Install shared library dependencies with npm ci --prefix packages/all-things-youtube before starting Wrangler. Configure OUTBOUND_PROXY_URLS or OUTBOUND_PROXY_URL in platform/.dev.vars for proxy fallback. Storyboards and exact frames still require Docker. Set YOUTUBE_EXTRACTION_BACKEND=container to use the rollback backend. See the repository document reference/engineering/WORKER_EXTRACTION.md for retry limits and rollout.