About LingoHub

LingoHub helps you learn the English people actually use in Australia and New Zealand — the everyday phrases, local expressions and workplace language that textbooks tend to miss.

How it works

LingoHub is a RAG (retrieval-augmented generation) app. Instead of letting an AI answer from memory, it first finds the passages in your learning materials that match your question, then asks the AI to answer using only those passages — and shows you which ones it used.

When a document is added

  1. 1

    Upload PDF

    POST /api/documents

  2. 2

    Extract text

    PdfPig reads each page

  3. 3

    Clean

    Drop page numbers, fix spacing and odd characters

  4. 4

    Chunk

    ~400 characters, 50 overlap

  5. 5

    Embed

    Voyage voyage-4 turns each chunk into 1024 numbers

  6. 6

    Store

    PostgreSQL + pgvector

When you ask a question

  1. 1

    Question

    Typed into the search box

  2. 2

    Embed question

    Same model as the chunks

  3. 3

    Vector search

    Cosine distance, top 5 chunks

  4. 4

    Build prompt

    Rules + 5 passages + question

  5. 5

    Generate

    Gemini writes the answer

  6. 6

    Answer

    With [1]-style citations and sources

The two flows meet in the database: the chunks stored in step 6 of the first flow are what the vector search looks through in the second. If none of the passages answer the question, LingoHub says so and asks before answering from the AI's general knowledge.

Design questions

  • A learner asks something that isn't in the documents. What happens, and why?

    Two layers. First, the prompt tells Gemini to answer only from the retrieved passages and, if they don't contain the answer, to start with one fixed sentence: “The provided document does not contain enough information to answer this question.” It must never guess.

    Second, the backend checks for that sentence and marks the answer as “not found”. The page then asks the learner whether they want an answer from the AI's general knowledge. Only after they say yes does the app send a second request. That answer skips retrieval, has no sources and is clearly labelled as not coming from the materials.

    Why: a grounded answer is the whole point of the app. A learner can't easily tell a plausible invented definition from a real one, especially for slang. A fixed sentence is something code can check reliably, unlike free text. Asking first keeps the learner in control, and the extra LLM call only happens when they want it.

    Known gap: retrieval always returns the top 5 chunks, even when none are close, so today the LLM decides whether they're relevant. The next step is a similarity threshold, tuned on an evaluation set, that skips the LLM call when nothing is close enough.

  • Why 400 characters with a 50-character overlap? Where do those numbers come from?

    From an evaluation, not a guess. The first version used 800 characters with a 150-character overlap, reasoned from the shape of the document. Then a 20-question test set (LingoHub.Eval) compared chunk sizes from 200 to 3,000 characters and overlaps from 0 to 150, using the real ingestion code.

    Size: each vocabulary entry is only about 70 characters, so an 800-character chunk holds about 11 different words, and its embedding is an average of all of them that matches none of them strongly. At 400 characters (about 5 entries) the right entry reached the top 5 for 95% of questions, up from 75%, and ranked first for 75%, up from 40%. Going smaller hurts again: at 200 characters (without overlap) 548 entries were cut in half and accuracy fell to 70%.

    Overlap: some is essential. At 400 characters with no overlap, 281 entries were split across two chunks (word in one, example sentence in the next). A 50-character overlap removed every split. Larger overlaps scored lower (85%), probably because near-duplicate chunks crowd the top 5.

    A bonus: the top 5 chunks are now about 1,900 characters per question instead of 3,900, so the retrieved context sent to Gemini is half the size. The caveat is the small test set: 20 questions means each one is 5%, so the direction is solid but the exact numbers are not.

  • Retrieval doesn't always find the right answer. How do you evaluate whether the RAG is good?

    Evaluate the two halves separately, because they fail differently.

    Retrieval (done): LingoHub.Eval has 20 questions, each labelled with the vocabulary entry that answers it, in four styles: English, Chinese, Chinese-to-English lookups and everyday scenarios. A question only counts as a hit if one retrieved chunk holds both the word and its example sentence. It reports hit rate@1 and @5 (is the right entry first / in the top 5?) and MRR (how high does it rank?). It reuses the real ingestion code, and a live mode sends the same questions to the running API; both give identical scores.

    What still misses: “TOIL” (time off in lieu) isn't found at all, and “arvo” and “technical debt” only come in at rank 5. Acronyms and short slang words carry little meaning for an embedding, which is exactly where keyword search is strong, so hybrid search (next question) is a likely fix.

    Generation (next): faithfulness (is every claim supported by the passage it cites?), correctness, citation accuracy, and refusals: does it say “not found” for out-of-scope questions without refusing in-scope ones? Review a small set by hand; use an LLM-as-judge to scale, and spot-check it.

    Also next: grow the set to 50+ questions, including deliberately out-of-scope ones, and rerun it whenever the chunk size, model, prompt or topK changes, like a regression test. Unit tests already guard the cleaning and chunking rules. In production, log the signals the API already returns (similarity scores, the not-found rate, tokens, latency) plus how often learners ask for general-knowledge answers.

  • Why embeddings and vector search instead of plain SQL or keyword search?

    Learners ask by meaning, not by the exact word. “What do Kiwis call the afternoon?” never contains “arvo”, so LIKE or keyword search finds nothing. An embedding places text with similar meaning close together, so the question lands near the right entry. It also works across languages: a question in Chinese still matches English entries.

    pgvector keeps the vectors in the same PostgreSQL as everything else. That means one database, the same transactions as the documents and chunks, and normal SQL filters alongside the distance (we only compare vectors made by the same embedding model).

    Trade-offs: keyword search is cheaper, easier to explain, and better at exact rare tokens such as names or codes. Embeddings cost one API call per question and can occasionally rank an exact match lower than a looser one.

    The strongest option is hybrid search: PostgreSQL full-text search plus vector search, with the two result lists merged (for example with reciprocal rank fusion). That's a natural next step.

  • What if the project grows from 600 chunks to 3 million? Does the architecture still work?

    The design holds (extract → chunk → embed → pgvector → LLM), but several parts need to change.

    Search: today there's no vector index, so every question computes its distance to every chunk. That's instant for 600 but not for 3 million: 3M × 1024 floats is about 12 GB of vectors to scan. The fix is an HNSW index in pgvector (approximate nearest neighbour), which brings queries back to milliseconds. Tune its recall/speed settings, and plan the memory: half-precision vectors (halfvec), or fewer dimensions if the model supports it, cut the size roughly in half or more.

    Ingestion: uploads currently chunk and embed during the HTTP request. At 3 million chunks (about 23,000 embedding calls at 128 per batch), that has to move to a background job queue with rate-limit handling and resumable progress. The existing “embed only missing chunks” endpoint is already a good basis for that.

    Quality: with many more documents, the top 5 can fill up with near-duplicates. Add metadata filters (per user or per document), hybrid search, and a reranking step after retrieval.

    Operations: connection pooling, read replicas, and partitioning by tenant if it's multi-user. 3 million vectors is comfortably within what pgvector handles with HNSW and enough RAM. A dedicated vector database only becomes worth it at much larger scale or with very high query rates.

Contact

Questions or feedback? hello@lingohub.example