rag-pet
LiveStage 2-2Retrieval-augmented generation over your own documents, with the retrieval shown instead of described.
- Chunk window
- 1,200 chars
- Embedding size
- 3,072 dims
- Ingest steps
- 7
- Query steps
- 6
- Python 3.13
- FastAPI
- pgvector
- React 19
- Vite
- Tailwind CSS 4
- Nx
- Fly.io
About
A FastAPI service does ingest and retrieval against Postgres and pgvector; a React console indexes sources, answers questions, and lists the exact chunks each answer was built from, so bad retrieval can be told apart from bad generation.
Highlights
- A Flow page that runs both pipelines live: the chunker with its overlap highlighted, a traced query down to the embedding, the exact prompt, and where the milliseconds went.
- Pydantic models are the only API contract: codegen turns FastAPI's OpenAPI schema into the frontend's types through the Nx graph, so a stale contract cannot compile.
- Every step links to the code that does it, and to a Learn article on how it fails.
How it flows
From start to goal, the path the data takes
Ingest
- Upload or paste
- Extract text
- Chunk
- Batched embed
- pgvector
Query
- Question
- Embed
- Nearest chunks
- Prompt
- Claude
- Answer + sources

