Get Started
All posts
Research

Türkiye Knowledge Base: More Accurate, More Current Answers

Language models answer from memory, and their knowledge freezes at a date. With the Türkiye Knowledge Base we aim to ground Simit's answers in verifiable sources.

Large language models are impressive but have two basic weaknesses: their knowledge freezes at a certain date, and they sometimes produce wrong information confidently. For topics such as legislation, public services or current economic data, that isn't acceptable.

Retrieval-augmented generation (RAG)

Our solution is for the model to retrieve relevant documents from a knowledge base before answering, and ground its answer in them:

  1. The question is turned into a semantic vector.
  2. The most relevant document chunks are found in PostgreSQL + pgvector.
  3. Those chunks are passed to the model as "context", with their sources.
  4. The model cites its source and never presents information as coming from the context when it doesn't.

What will it include?

  • Openly licensed public documents and official announcements
  • Documents users upload themselves (for their own use only)
  • User-specific memory, with permission
  • Later, controlled web search

Status

The retrieval layer is already part of Simit GPT's architecture; for now it returns an empty context. The knowledge base arrives with Phase 2.