Back to the library

Build Your Own Kindle Brain

guide

Note: Paste this whole guide into Claude Code or Codex and go. Each step is one job and one prompt. Do them in order, and check each one works before you start the next.

You highlight a passage because it hit you. Then it goes into a list on Amazon you never open again.

This guide builds a private app on your Mac that holds every Kindle highlight and note you've ever made, and lets you search them by topic, even when the passage never uses the word you typed. Mine has 263 books and 14,183 highlights, and a search comes back in under half a second.

Nothing leaves your Mac. The search model runs on your own machine and there's no account to make.


What you're building

  1. An exporter that pulls every highlight and note out of your Kindle notebook in one run
  2. A library on your Mac that never loses a passage, however many times you re-import
  3. Topic search, using a small model that runs locally
  4. The app: search, a wall of your book covers, a shelf where the books you marked up most have the thickest spines, saved Ideas, and old highlights brought back on the home screen
  5. An Update from Kindle button that only pulls in what's new
  6. A launcher, so it opens with one click

What you need

  • A Mac. Apple Silicon is best (I built mine on a MacBook Air, M4, 16 GB).
  • Claude Code or Codex
  • Chrome, or any browser with a developer console
  • About 1 GB free for the search model
  • Highlights in your Kindle notebook at read.amazon.com/notebook

1. Start the project

Make an empty folder, open Claude Code or Codex in it, and paste this.

I want to build Kindle Brain: a private app on my Mac that lets me search every Kindle highlight and note I've ever made by topic. Rules for the whole project: everything stays on this Mac (no cloud AI, no accounts, no analytics), every server binds to 127.0.0.1 only, and my data lives in a data/ folder that is gitignored and never committed. Use Python and SQLite (with FTS5) for the backend, a small React app built with Vite for the interface, and Ollama for the local model. We'll build it one step at a time and I'll give you each step. For now, set up the project folder, the .gitignore, and a README that explains how to run it. Ask me before installing anything system-wide.

2. Export every highlight from Amazon

Tools like Bookcision export one book at a time. With a few hundred books that's an afternoon of clicking. Have your AI write an exporter that does the whole notebook in one run, from inside your own signed-in browser tab.

Write scripts/export-notebook.js, a script I paste into the browser console while I'm signed in at read.amazon.com/notebook. It only reads from Amazon. It should:
1. Find every book by paging through /notebook?library=list&token=… until there's no next-page token. Each book is a .kp-notebook-library-each-book element whose id is the ASIN, and the next token is in .kp-notebook-library-next-page-start.
2. For each book, page through /notebook?asin=…&token=…&contentLimitState=… (next token in .kp-notebook-annotations-next-page-start) and collect every highlight and note with its location, its page number when there is one, and Amazon's annotation id.
3. Save a note that has no highlight as its own standalone note, not as an empty highlight.
4. Read each book's "N Highlights | N Notes" header so we can check counts later, and each book's last-annotated date.
5. Use same-origin fetch with my existing session. Retry 429s and 5xx errors with backoff, pause briefly between requests, and stop if a page token ever repeats.
6. Keep the raw HTML of every page in the export, so we can re-parse it later without asking Amazon again.
7. Show a small progress panel with Save progress and Stop buttons, and download one JSON file at the end.

To run it, open read.amazon.com/notebook in Chrome, press ⌥⌘J to open the console, paste the script and press Enter. The first time, Chrome makes you type allow pasting before it lets you paste anything.

What I found in my export:

  • Amazon gives no date for individual highlights. Each book only has a "last annotated" day. Don't let your AI fill in highlight dates from the import time.
  • Two different highlights can sit at the same location (I had five of them), so location alone can't identify a highlight.
  • Page numbers only exist for some books. About half of my highlights have one.
  • The export has highlights and notes. Kindle bookmarks aren't in it.

3. Build a library that never loses anything

Build library.py on SQLite. Tables for books (ASIN as the id, title, authors, and the last-annotated date exactly as Amazon gives it), annotations (text, my note, a standalone-note flag, location, page, and a kindle:// link back to that spot in the book), the raw imports, and an FTS5 index for keyword search. Rules:
1. Re-importing the same file adds nothing.
2. Imports never delete. A passage missing from a newer file stays.
3. When there's no stable id, fingerprint each record on book + text + location + record type, because two highlights can share a location.
4. Keep my notes in their own field, never mixed into the book's text.
5. Leave a date empty when Amazon doesn't give one. Never use the import time as a highlight date.
Add scripts/import_archive.py to import an export from the command line, and tests for re-importing, standalone notes, and two highlights at the same location.

Then import the file the exporter saved to your Downloads folder.

4. Add topic search

You type a question like "what have I read about committing to one thing?" and get the passages that answer it, whatever words the author used.

Add local topic search. Run Ollama as a service the app owns, on 127.0.0.1:11439, with OLLAMA_NO_CLOUD=1 and its models stored in data/models, and pull the embeddinggemma model.
Embed every passage, with my note appended, using EmbeddingGemma's retrieval prefixes: passages as "title: <book title> | text: <passage>" and searches as "task: search result | query: <search>". Split passages over 1,200 characters into chunks with 120 characters of overlap, and never let the model silently cut text off.
Store each vector with a hash of its text and the model's digest, so indexing can stop and resume and skips anything unchanged.
At search time, score every vector (it's fast at this size), merge chunks back into one result per passage, blend in the keyword results, drop exact duplicates, give each extra result from the same book a small penalty so one book can't fill the list, and return the top 15. If the model isn't running, show keyword results and say so.
On Apple Silicon, start Ollama with arch -arm64.

That last line fixes the bug that cost me the most time. If your project's Python is the Intel build, it runs under Rosetta, and anything it launches runs as Intel too. Ollama loses the GPU and search keeps falling back to plain keyword results. Check with:

.venv/bin/python -c "import platform; print(platform.machine())"

If it prints x86_64 on an M-series Mac, you need the arch -arm64 fix. After it, my searches came back in under half a second. The first search after opening the app takes a second or two while the model loads.

Skip the AI summary, at least at first. I built a "find patterns" feature on a small local model (Qwen3 4B) that summarized search results. It was slow, a new search had to wait while it ran, and it wasn't finding real patterns. Topic search is the part I actually use.

5. Build the app

Build the interface as a React app served by a small Python server (server.py) on 127.0.0.1:4180.
Explore: one search box. Results are passage cards with the book, author, page and location. Show my note separately from the book's words, and add a My notes filter.
Library: a wall of book covers, plus a shelf view where each spine's width scales with how many highlights that book has. Sort A–Z, Recent, or Most highlighted.
Book page: every passage from that book.
Saved Ideas: pick passages from any search, save them as an Idea with a title and my own thought, add more passages to an Idea later, and copy a whole Idea with citations.
Home: bring back a few old highlights to reread, with a button to shuffle them.
Protect the server: check the Host and Origin headers, require a custom header on every write, no CORS, and don't log searches.

For the covers:

Fetch book covers by ASIN from https://images-na.ssl-images-amazon.com/images/P/<ASIN>.01.LZZZZZZZ.jpg into data/covers. Amazon answers an unknown ASIN with a tiny transparent GIF instead of a 404, so only keep files that start with the JPEG header and are bigger than 1 KB.

Covers by ASIN match the exact edition you read. Every book in my library got one.

6. Add Update from Kindle

Add an Update from Kindle button.
Step one: the app hands me a copy of export-notebook.js with every book's last-annotated date baked in, so the script only downloads books that are new or changed. Amazon only dates books by the day, so also re-read anything annotated today or yesterday. If nothing changed, say "Already up to date" and download nothing.
Step two: the app finds the newest export in ~/Downloads, keeps a copy in data/imports, imports it without deleting anything, fetches covers for new books, and indexes only the new passages. If macOS blocks access to Downloads, tell me to use an Import JSON button instead.

Deleting a highlight on your Kindle doesn't remove it here. Imports only ever add.

7. Make it open with one click

Make a small Mac app bundle, Kindle Brain.app, that runs launch.py. The launcher starts Ollama and the server if they aren't running, reuses them if they are, and opens the browser. Use a lock file so double-clicking twice can't start two copies. Add scripts/stop.py to shut everything down, a Download backup button that zips the database, the original exports and a manifest, and a restore script that only ever restores into a new folder.

The app bundle finds your project by looking next to itself, so leave it in the project folder and put an alias in your Dock. A copy of the bundle won't find anything.

Get the next guide when it's ready.

You're on the list.