Skip to content
Lecture Revision Lab
All decision records

Decision record DR-001 · Accepted · decided 9 October 2026

Browser-only architecture with bring-your-own-key AI and a local audit log

Decision: Run the whole lab in the visitor's browser, make every AI feature optional and powered only by the visitor's own API key called directly from the browser, and record every model call with the human decision on its output in an audit log kept in that browser.

StatusDecidedRecordedOwner
Accepted9 October 202610 October 2026Sunchuangyu (Rin) Huang

Context

The 2024 prototype was a Python script that read an OpenAI key from a local .env file and stored lecture notes in SQLite. Reviving it as a public website raised four problems. There is no budget for model calls, so the site cannot pay for visitors' generations. A server route holding a site key would be a public endpoint that anyone could abuse. Visitors should not have to trust this site with their own key. And the notes people want to revise from are often private or copyrighted, so they should not travel to a server this site controls.

The retrieval half of the prototype needs no model API at all: the embedding model is small enough to run in a browser, and the demo textbook's vectors can be computed once and shipped as static files.

Decision

  • The site is static: Next.js pages with no server code that handles data. Uploaded files are parsed, chunked, embedded and stored in the visitor's browser (IndexedDB). Retrieval and the retrieval evaluation work without any key.
  • AI generation is optional. The visitor opens "AI settings", picks Anthropic (the default) or OpenAI, and pastes their own key. Keys are kept per provider in sessionStorage, or in localStorage only if the visitor ticks "remember on this device", and a "forget keys" button removes them.
  • The default model is Claude Haiku 4.5, the cheapest tier, with Claude Sonnet 5.5 and Claude Opus 5.5 as options. The OpenAI model id is editable and defaults to gpt-4o-mini, the prototype's model.
  • Requests go straight from the browser to api.anthropic.com (through the official SDK, which sends the anthropic-dangerous-direct-browser-access header) or api.openai.com. The key travels only in request headers to the provider.
  • Sonnet 5.5 and Opus 5.5 requests on Generate and Ask opt in to Anthropic's server-side refusal fallback. The model that actually answered is read from the response and recorded, so a fallback answer is never attributed to the model that was asked. Generation evaluation runs do not opt in, so a refusal there counts against the model being evaluated.
  • Every call, successful or not, is appended to the ai_log table with its feature, provider, model, input (system prompt, any earlier conversation turns sent with it, user prompt and passages), output, latency, token usage, sampling temperature and the human decision on the output (pending, partial, accepted, edited, rejected or mixed). Anything shaped like a provider key is redacted before a row is written. The log is viewable at /ai-log and exportable as JSON or CSV.
  • Every AI output is labelled "AI-generated" where it appears and in revision exports, and pre-recorded examples are labelled as recordings. Revision exports contain only items a person accepted; the audit log and evaluation exports keep every output so the record can be audited.

Options considered

OptionWhy not
A server route with the site's own keyCosts money for every visitor and turns the site into a public endpoint that invites abuse.
A server route that forwards the visitor's keyThe key would pass through a server this site controls, which visitors would have to trust.
Only pre-recorded outputs, no live callsHonest and free, but visitors could not try their own topics or repeat the generation evaluation.
No AI features at allLeaves the prototype's purpose, drafting revision material, unfinished.
Browser-direct, bring your own key (chosen)No cost to the site, no key or notes on a server, and anyone can repeat a run with their own model choice.

Why

It keeps the governance simple to state and easy to check in the code. No server of this site ever receives a secret, the visitor sees exactly what is sent and where it goes, every output is labelled, and a person decides what happens to it. The design is informed by the Australian Government's policy for the responsible use of AI in government, the transparency principles of the EU AI Act and the NIST AI Risk Management Framework. It does not claim compliance with any of them.

What happened

  • The first revival build (9 October 2026) shipped a single-key dialog with Claude Opus 5.5 as the default. On 10 October the default moved to Claude Haiku 4.5, keys became per provider, the old storage format is migrated on first read, and the audit log gained the feature name, the human decision and the fallback flag. IndexedDB moved to version 2 for the evaluation runs table; rows written by version 1 simply lack the new fields.
  • A typical Generate call on the demo text with five passages sends about 1,350 tokens and receives about 650 (estimated at four characters per token from the pre-recorded examples), which is roughly half a US cent on Haiku 4.5 at list price, and about four cents for the eight-topic generation evaluation. These are estimates; the provider bills the visitor.
  • The first build exported accepted and not-yet-reviewed items, which contradicted "nothing is exported until a person has reviewed it". Since 10 October revision exports contain accepted items only, and the export buttons stay disabled until something is accepted.
  • Review before merge found three gaps, fixed on 10 October. The Ask log recorded the latest prompt but not the earlier turns sent with it, so the log could not show everything that left the browser; it now stores them. Leaving the generation evaluation mid-run kept calling the provider on the visitor's key; leaving the page now cancels the call in flight and marks the run stopped. And a tab still on the old version could block the storage upgrade and hang a paid call on "Generating…"; the page now asks the visitor to close the other tab, and local logging can no longer hold up a result.
  • The site publishes no live model outputs and no generation scores, because there is no budget to produce them. The pre-recorded examples were written once by Claude Opus 5.5 during the rebuild and are labelled as recordings.
  • Tests check that the key never appears in a request body, an error message or a log row, and that a failed call is still logged with its partial output. Review before merge found that a refused, cut-off or invalid answer was logged without the token usage the provider reported, although it was billed; failed calls now keep the usage, the model that answered and the latency.
  • A key in browser storage is readable by any script running on this origin. On 10 October main added a production Content-Security-Policy (web/next.config.ts): scripts only from this site (plus the ONNX runtime from jsDelivr), and connect-src limited to this site, the two providers and the model hosts, so an injected script could not send a stored key elsewhere. It still allows 'unsafe-inline' scripts, which Next.js needs for the inline bootstrap of static pages.

What I'd change

  • Move the Content-Security-Policy to nonce-based scripts so 'unsafe-inline' can be dropped, which would need the pages rendered per request.
  • Offer a per-session spending cap that counts estimated tokens before each call.
  • With a small budget, commit one labelled reference run of the generation evaluation (model, date, settings and exported JSON) so the comparison is visible without a key.

Source: docs/decisions/DR-001-browser-only-byok-architecture.md in the project repository (GenAI-Lec-Gen, private).