AI / SaaS
DocQuery
Upload a PDF, ask it questions, get answers grounded in the document.
- Year
- 2024
- Role
- Sole engineer — architecture, frontend, backend
- Type
- AI / SaaS
What it is
DocQuery is a document question-answering product. You upload a PDF, it gets parsed, chunked, embedded and stored in a vector index, and then you can ask questions about it in a chat interface. Answers stream back grounded in the pages that actually matter. There is a free tier with page and file-size limits, and a pro plan behind Stripe.
By the numbers
- 7
- External services wired
- 0
- Hand-written API types
- 2
- Embedding providers compared
- 01UploadUploadThing
- 02Parsepdf-parse
- 03Chunkoverlap window
- 04EmbedOpenAI / Cohere
- 05IndexPinecone
- 06Retrievetop-k + context
- 07StreamAI SDK
The problem worth solving
The interesting part of a RAG product is not the chat interface — it is retrieval quality. A naive chunk-and-embed pipeline will confidently answer a question using text from the wrong page, and the user has no way to tell. The failure is silent, which makes it the worst kind.
Most of the engineering went into the boundary between the document and the model: how big a chunk should be, how much chunks should overlap so a sentence split across a boundary is not lost, how many neighbours to retrieve, and how much of that retrieved context to actually hand the model before it starts ignoring the middle of it.
What I built
Uploads land through UploadThing so large PDFs never pass through a serverless function body. The file is parsed, split into overlapping windows, embedded, and upserted into Pinecone under a namespace scoped to the file — so retrieval can never leak across documents or across users.
At query time the question is embedded with the same model that indexed the document, the nearest chunks are pulled back, and the question plus retrieved context plus the recent turns of the conversation are streamed to the model through the Vercel AI SDK. The response streams token by token into the client rather than making the user wait on a complete answer.
Billing is a Stripe subscription with a free tier enforced server-side at upload time. Auth is Kinde, so I did not hand-roll sessions.
The part I would show a reviewer
The whole stack is typed end to end with no generated client and no hand-written request types. Prisma types the database, tRPC v11 infers the router's input and output types straight into the React client, and Zod validates at the boundary. Renaming a database column surfaces as a red squiggle in a component, at author time, not as a runtime error in production.
That property is the reason I would build it this way again. It removes an entire category of bug rather than catching it later.
Decisions
Four calls I would make the same way again
01
Two embedding providers, not one
I wired both OpenAI and Cohere so retrieval quality could be compared on the same corpus instead of assumed. Different providers genuinely disagree about which chunk is nearest, and on a document-QA product that disagreement is the product.
02
Namespace per document
Pinecone namespaces scope every query to a single file. A retrieval bug can therefore return a wrong chunk, but never someone else's chunk. Isolation enforced by the data layer beats isolation enforced by remembering to add a filter.
03
Stream, don't wait
A grounded answer can take several seconds to generate. Streaming turns that into visible progress instead of a spinner, which changes how fast the product feels far more than any optimisation of the actual latency.
04
Enforce plan limits on the server
Free-tier page and size limits are checked before ingestion, not in the UI. The client-side check is a courtesy; the server-side check is the rule.
Stack
Retrieval
- Pinecone
- LangChain
- OpenAI embeddings
- Cohere embeddings
- pdf-parse
Data
- PostgreSQL
- Prisma
- Zod
Transport
- tRPC v11
- TanStack Query
- Vercel AI SDK
Platform
- Next.js 14
- Kinde Auth
- Stripe
- UploadThing
Next case study
Commerce Platform