rag-pet

LiveStage 2-2

Retrieval-augmented generation over your own documents, with the retrieval shown instead of described.

Fig 1 / 4Flow: both paths, every step linked to the code that runs it
Chunk window
1,200 chars
Embedding size
3,072 dims
Ingest steps
7
Query steps
6
  • Python 3.13
  • FastAPI
  • pgvector
  • React 19
  • Vite
  • Tailwind CSS 4
  • Nx
  • Fly.io

About

A FastAPI service does ingest and retrieval against Postgres and pgvector; a React console indexes sources, answers questions, and lists the exact chunks each answer was built from, so bad retrieval can be told apart from bad generation.

Highlights

  • A Flow page that runs both pipelines live: the chunker with its overlap highlighted, a traced query down to the embedding, the exact prompt, and where the milliseconds went.
  • Pydantic models are the only API contract: codegen turns FastAPI's OpenAPI schema into the frontend's types through the Nx graph, so a stale contract cannot compile.
  • Every step links to the code that does it, and to a Learn article on how it fails.

How it flows

From start to goal, the path the data takes

Ingest

  1. Upload or paste
  2. Extract text
  3. Chunk
  4. Batched embed
  5. pgvector

Query

  1. Question
  2. Embed
  3. Nearest chunks
  4. Prompt
  5. Claude
  6. Answer + sources