Bogart the Taskmaster

InternalStage 1-2

A desktop QA agent: Claude writes the test plan, a fast decision model runs it in your own browser.

Fig 1 / 7Plan stage: the ticket on the left, Claude's test plan on the right
Decision latency
~300 ms
Pass threshold
p_yes ≥ 0.5
Sprite pipeline
~1,400 lines
Lore
8 chapters

Internal system at Boost Capital, not public.

  • Electron
  • Node.js
  • Claude Code CLI
  • Playwright / CDP
  • PostgreSQL
  • WebGL
  • Canvas

About

An Electron app for the step between "fixed" and "sent to QA". It reads the ticket, its comments, the commits and PRs tagged with it and earlier runs, then drives each step over the Chrome DevTools Protocol and gives every check a verdict, a probability and a screenshot.

Highlights

  • Ticket in, plan out: claude -p with a JSON schema writes a structured, editable plan; Jev, TypeSafe's decision model, runs each step in the signed-in browser.
  • Secrets stay on the machine: saved logins are encrypted with Electron safeStorage, and the models only ever see placeholders.
  • Safe on real accounts: runs hand off to the person for 2FA, SSO and irreversible actions, and a failed run continues from step N.
2 more
  • Full history in Postgres: content-hashed plan revisions with step diffs, per-step evidence, and a report built without a model call.
  • Bogart himself is drawn in code: a 32×32 sprite pipeline, a WebGL pixel-fire effect and an eight-chapter manga.

How it flows

From start to goal, the path the data takes

  1. Ticket, comments, linked PRs
  2. claude -p · JSON-schema plan
  3. Editable steps
  4. Jev drives the browser
  5. Verdict + evidence per step
  6. Report · Postgres history

Shown with a sample test against the public demo shop saucedemo.com, not a real ticket.