Why this exists
I have a shelf of books I bought and never finished. I have a laptop full of PDFs I downloaded and never read. The summaries on Blinkist are too shallow. ChatGPT will talk about any book confidently, but it's making half of it up — it's drawing on reviews and training data, not the actual text in front of me.
I wanted something simpler. Drop a PDF in. Ask a question. Get an answer that's grounded in that specific book, with the passage it came from. No hallucinated chapter numbers, no generic life advice borrowed from somewhere else.
I built it for myself one weekend. Then I kept opening it. So now it's a project.
How it works
The flow is boring on purpose:
- Upload a PDF.
- The book gets chunked, embedded, and stored.
- You ask a question. The relevant chunks come back. The model answers using only those chunks, and shows you the passage it leaned on.
That's it. No fancy agents. No tool calls. Just retrieval and a tight system prompt that tells the model to refuse politely when the book doesn't actually cover what you asked.
An example
From a session on Atomic Habits:
Me: What does Clear mean by "habit stacking"? Give me one of his actual examples.
App: Habit stacking is the practice of attaching a new habit to an existing one. The example Clear gives in chapter 5: "After I pour my cup of coffee each morning, I will meditate for one minute." The existing habit (coffee) becomes the trigger for the new one (meditation).
The passage it pulled from is shown right next to the answer. If the book doesn't have it, the app says so instead of inventing something plausible.
What's in it
- PDF in, conversation out. No format gymnastics. Drop the file, start asking.
- Citations on every answer. The chunk it used is visible. You can verify or jump straight to the source page.
- Refusal when grounded. If the book doesn't say it, the app doesn't either.
- Session memory per book. You can pick up where you left off and the app remembers what you've already discussed.
- Local first. The PDF stays on your machine. Nothing gets uploaded to a vendor library.
What I'd compare it to
| This | NotebookLM | Blinkist | |
|---|---|---|---|
| Source | Your PDF | Your upload | Their catalog |
| Format | Chat, with citations | Chat plus podcast | Pre-recorded summary |
| Depth control | Ask whatever you want | Fixed | 15 minutes, fixed |
| Grounding | Refuses when unsure | Generally good | Not applicable |
NotebookLM is the closest comparison and it's genuinely good. The reason I still use mine: the citations are tighter, the refusal behaviour is stricter, and I trust it on a single book in a way I haven't quite trusted NotebookLM yet.
Tech stack
- PDF parsing — chunked by section heading, with overlap.
- Embeddings — stored locally in a small vector index.
- Retrieval — top-k with a rerank pass before it hits the model.
- Model — Claude for the answer step. The system prompt does most of the heavy lifting.
- Frontend — plain HTML and a bit of JS. No framework.
Where this is going
Still a weekend project. Still rough around the edges. The next things on the list:
- Better handling of books with diagrams and tables — right now they get flattened into messy text.
- Voice mode, eventually. Reading on a screen is fine; talking to a book on a walk is the real unlock.
- A small library of pre-processed public-domain books so first-time users have something to play with before they upload anything.
Not promising any of that. It exists, I use it, and that's enough for now.