AI & Automation

Nexus: I Built a Self-Hosted Knowledge Base That Cites Its Sources

Jabane Mohamed Ayoub11 min read

Nexus title card on a pink-to-purple gradient with a knowledge-graph illustration, reading: a self-hosted knowledge base that cites its sources
Table of contents
  1. What it includes
  2. Why this project exists
  3. Not just Slack
  4. How it’s built
  5. Slack is the hard part
  6. Retrieve, fuse, rerank
  7. Two ways in
  8. Quick start
  9. Lessons from standing it up on Windows
  10. Who this is for
  11. Further reading
  12. Related resources
  13. FAQ

Nexus is a self-hosted knowledge base that answers questions from everything a team already writes down: Slack threads, code, wikis, tickets and web documentation. Every answer comes with citations back to the original source. It runs on your own infrastructure, and it can run with no cloud model at all.

I built it because half of my week in marketplace and channel integration work went into questions that had already been answered once, somewhere: which channel sign belongs to which marketplace, whether a wiki page is still accurate, why a feed failed with an error I’d definitely seen before. The answer always existed. It just lived in a Slack thread, a help-center page and a Jira comment that had never met each other.

It’s still in beta, and the repository is private while I finish stability and security work. See Quick start if you want early access.

What it includes

  • Ten connector types: Slack, GitHub, Confluence/wiki, Notion, Google Drive, Jira, Linear, web pages (via trafilatura and crawl4ai), file uploads, and custom Python connectors for anything else.
  • Hybrid search: semantic, full-text and recency signals fused with reciprocal rank fusion, then reranked against the actual question.
  • Three interfaces: a web UI with streamed, cited answers; a CLI; and an MCP server so Claude Code, Cursor or any MCP client can query it directly.
  • Provider-agnostic models: Anthropic, OpenAI, OpenRouter, Ollama and LM Studio, with a different model per capability, down to fully local on one laptop.
  • Projects and roles: sources are scoped into named projects, with admin/member/viewer access and per-project API keys.

Why this project exists

Ask anyone who has joined a team recently what they spent their first month doing, and you get the same three questions: “Where do I find X?”, “Who knows about Y?”, “Is Z still true?” The answers are never missing. They’re scattered across tools that each have their own search box and no visibility into each other. A channel ID lives in a Slack thread, the process around it sits in a help-center page, and the reason it changed is a comment on a Jira ticket that got resolved eight months ago.

The tempting fix is to declare one system the source of truth and move everything into it. In practice this never survives contact with how people actually work: a suggested edit belongs in a document, a discussion belongs in Slack, a decision belongs in a pull-request review, and forcing any of those into a wiki page makes both worse.

So Nexus does the opposite. It reads from each system where it already lives, indexes it, and gets out of the way. Each source is one connector, and nobody has to change how they work for the tool to see their work.

Not just Slack

Most of this post focuses on Slack because it’s the hardest source to search well, but Slack is one connector among ten, not the main event. Google Drive makes the business case better than Slack does: contracts, pricing sheets, partner onboarding decks and product specs pile up in Drive folders that only the person who created them can really navigate. Point the Drive connector at them and those documents become searchable and citable right next to Slack threads, code and tickets, no migration, no asking anyone to re-file a single file.

For a business, the payoff isn’t the search box itself. It’s what stops happening: fewer weeks spent re-explaining onboarding to new hires, fewer support answers given from memory instead of the policy that’s actually current, and less knowledge that walks out the door when someone changes teams, because it was never only in their head, it was in a Drive folder Nexus had already indexed.

How it’s built

Nexus is three things stacked on top of each other: a platform for collecting and storing internal data, a platform for querying it, and a layer for auth, auditing and cost control on top.

flowchart TB
  accTitle: Nexus architecture
  accDescr: A Next.js frontend and FastAPI backend exchange answers over SSE. The backend queries Postgres with pgvector, and fans out to a Redis cache, a sync scheduler and four LLM provider roles. The scheduler pulls incrementally from Slack, GitHub, Confluence, Notion, Google Drive, Jira, Linear, the web and custom connectors.

  FE[Next.js frontend]
  BE[FastAPI backend]
  DB[(Postgres 17 + pgvector<br/>HNSW vectors &middot; GIN full-text)]
  REDIS[(Redis cache)]
  SYNC[Sync scheduler<br/>incremental &middot; idempotent]
  LLM[LLM providers<br/>embed &middot; distill &middot; rerank &middot; synthesize]

  FE <-->|SSE| BE
  BE --> DB
  BE --> REDIS
  BE --> SYNC
  BE --> LLM

  subgraph Sources["Sources, read incrementally by connector"]
    direction LR
    SLACK[Slack]
    GH[GitHub]
    WIKI[Confluence / Wiki]
    NOTION[Notion]
    DRIVE[Google Drive]
    JIRA[Jira]
    LINEAR[Linear]
    WEB[Web: trafilatura &#183; crawl4ai]
    CUSTOM[Custom connector]
  end

  SYNC --> Sources

Postgres with pgvector sits at the centre. I looked at a dedicated vector database first, but the vectors need to sit next to the metadata they get filtered by (project, source, permissions), full-text search is already built into Postgres, and backups, migrations and access control stay one system instead of three.

Every source type gets its own embeddings table (slack_embeddings, code_embeddings, wiki_embeddings, custom_embeddings) with an identical schema, exposed to search as a single view: embeddings_unified. Each source can be tuned on its own, and the query side only ever sees one table.

The connector interface stays deliberately small:

configure()          # how to reach the source
fetch(watermark)      # what changed since last time
transform()           # shape it into embedding rows

Syncs are incremental and idempotent: each chunk gets a SHA-256 hash, so unchanged content skips its embedding call entirely, changed content is replaced in place, and content deleted at the source is cleaned up on the next pass.

Slack is the hard part

Slack carries the most current engineering discussion on any team, and it’s also the source that breaks naive RAG the fastest. Embedding raw messages sounded fine until I actually tried it:

  • Information density varies wildly. “yeah sure, thanks” and a detailed explanation of a race condition are both one Slack message.
  • Short messages often win on cosine similarity over the long message that actually answers the question.
  • A message’s meaning depends on the thread around it, not the message alone.

So Nexus treats the thread as the unit of knowledge, not the message. The connector re-fetches the full conversation (parent and every reply) from a stored watermark, so the stored row always reflects the whole discussion, not a stale fragment of it.

Before embedding, a small, fast model (the distill capability) reads the entire thread and extracts four things:

QUESTION:    <a one-line question an engineer would actually search for>
SUMMARY:     <2-3 sentences>
RESOLUTION:  <the answer, if there was one>
REFERENCES:  <systems and code identifiers mentioned>
flowchart LR
  accTitle: Distilling a Slack thread
  accDescr: The full thread, parent and every reply, is re-fetched on every update and passed to a distill model, which extracts a one-line question, a summary, the resolution and referenced systems. That normalized document is embedded and stored as one row, while the raw thread is kept alongside it for full-text search.

  T[Full thread<br/>parent + every reply]
  D{{Distill model}}
  A[Question &middot; summary<br/>resolution &middot; references]
  E[(embeddings_unified row)]
  R[(Raw thread<br/>kept for full-text search)]

  T --> D --> A --> E
  T -.also kept as.-> R

That normalized document is what gets embedded, not the transcript. The raw thread stays alongside it, keyword-searchable through a Postgres full-text index, and the extracted references give search an extra, exact-match signal.

No single retrieval technique gets trusted alone. Four run for every query and cover each other’s blind spots:

Technique What it catches
Full-text search Exact tokens embeddings blur: error strings, flag names, host names
Embedding search Paraphrase: different words, same problem
Inverse document frequency Signal vs. filler, so “sounds good, thanks!” doesn’t outrank a rare config flag
Age decay Slack answers expire; the newer of two equally relevant threads wins

Cerebras describes a similar system for their own internal knowledge base, and adds one technique Nexus doesn’t have yet: bursting, embedding individual same-author message runs inside a thread, scored by token rarity, length and reactions, so a single buried message can be found on its own even when the thread-level summary misses it. It’s a sensible next step; today thread-level distillation is Nexus’s only Slack embedding.

Retrieve, fuse, rerank

For each query, Nexus can run up to seven retrievers in parallel (vector, full-text, code, wiki, Slack keyword, thread-summary, and an optional entity graph), chosen automatically from the source types present in the project, so a project with no code never pays for a code search.

flowchart LR
  accTitle: Retrieve, fuse, rerank
  accDescr: A query fans out to up to seven retrievers in parallel, chosen automatically from the source types present in the project. Their ranked lists are combined with reciprocal rank fusion, reranked against the original question, and expanded with neighbouring chunks before reaching the answer.

  Q[Query]
  V[vector]
  FT[fulltext]
  C[code]
  W[wiki_vector]
  SF[slack_fts]
  TS[thread_summary]
  G[graph, optional]
  RRF{{Reciprocal rank fusion}}
  RR[Rerank 0-10<br/>4s time limit]
  CTX[Add neighbouring chunks]
  ANS[Answer / MCP result]

  Q --> V --> RRF
  Q --> FT --> RRF
  Q --> C --> RRF
  Q --> W --> RRF
  Q --> SF --> RRF
  Q --> TS --> RRF
  Q --> G --> RRF
  RRF --> RR --> CTX --> ANS

Their ranked lists aren’t comparable on raw scores: a cosine distance and a text-search score live on different scales. I combine them with reciprocal rank fusion: for every document, add weight / (60 + rank) for each list it shows up in. That constant means consensus across retrievers beats one strong vote from a single one. After fusion, age decay is applied, duplicate chunks are merged back to one source, and a cap on how many results one source can contribute keeps the top of the list diverse.

The fused list still gets one more pass: a small reranker model scores each candidate against the original query from 0 to 10. Scores are cached per (model, query, document), and reranking has a hard four-second time limit. If it’s slow, search falls back to the fused order instead of blocking. Once the final ranking is set, neighbouring chunks get pulled back in, so a matched wiki paragraph shows up with the heading and caveats that chunking had split away from it.

Cerebras runs an LLM planning pass before any of this, to decide which tools and sources are worth calling for a given question. Nexus skips that step: retrievers are chosen deterministically from which source types exist in the active project, so there’s no extra model call and no added latency before search even starts. The trade-off is less nuance (Nexus can’t decide mid-query that a question probably doesn’t need the graph retriever), but every query’s cost and latency stay predictable.

Two ways in

MCP. Retrieval is exposed as small, LLM-free tools: search, search_slack, search_code, search_wiki, who_knows, recent_prs, ripgrep_code, cheap and fast to call. An agent like Claude Code or Cursor does the orchestration itself: which tool, in what order, what to do with the result.

Web UI. The same retrieval runs as a full pipeline: search, evidence shaping, then a synthesis model that writes the answer, streamed token by token over SSE, with numbered citation chips that open the source. From the outside it’s “ask a question, get an answer.” Underneath, it’s the same retrieval an MCP client could call step by step.

Quick start

Nexus is in beta, and the repository (SilentJMA/DieKG) is private for now while I finish stability and security work. Get in touch if you want early access. Once you’re added as a collaborator, setup is one script:

git clone https://github.com/SilentJMA/DieKG.git && cd DieKG
./setup.sh          # build, start, migrate, bootstrap admin
# open http://localhost:3000 -> Create Admin Account
# add an LLM key to .env; set provider routing in config.yaml

./setup.sh --prod adds Caddy in front with automatic HTTPS for a real deployment.

Lessons from standing it up on Windows

I recently brought up a fresh instance on Windows, driving the embed and synthesis models through LM Studio instead of a hosted API. What actually broke:

  1. Match the vector dimension to the embed model. The default column assumes 1536 dimensions (OpenAI). A local nomic-embed-text model gives 768. The column type, the HNSW index and the config all have to agree, or inserts fail silently different ways depending on which one is wrong.
  2. Reloaded settings don’t always reach every long-lived component. Saving provider settings rebuilt the API’s provider registry, but the background sync scheduler kept the old one in memory, so scheduled syncs kept using the previous provider until I restarted the process.
  3. A strict CSP blocks the browser from reaching the API silently. With the UI on one port and the API on another, connect-src 'self' just fails every request with no useful error. The UI reports “cannot reach the backend” and gives you nothing to go on.
  4. Confirm a restart actually killed the old process. On Windows, orphaned workers can keep listening on the port next to the new one, so requests get randomly split between old and new code until you notice.
  5. Local-first genuinely works. All four model capabilities (embed, distill, rerank, synthesize) ran on one machine. Answer quality tracked the synthesis model far more than any other setting.

Who this is for

Any team whose knowledge is split across a help center, a wiki, a chat tool, tickets and a few repositories, and who wants one question box with citations instead of five search boxes with none. It’s a good fit if you want control over where the model runs (including fully offline), and a bad fit if you need a vendor to own security review, uptime and support on day one.

Further reading

Nexus’s retrieval design leans heavily on Cerebras’ public write-up about building their own internal knowledge base, worth reading in full, especially their sections on bursting and LLM-based query planning, which Nexus does differently.

The specific techniques above come from:

FAQ

Is the code public?

Not yet. Nexus is in private beta while I finish stability and security work, so the SilentJMA/DieKG repository isn’t public. Get in touch if you want early access as a collaborator.

Does Nexus only work with Slack?

No — Slack just gets the most detail in this post because it’s the hardest source to search well. Nexus ships ten connector types, including Google Drive, GitHub, Confluence, Notion, Jira and Linear, and every one is indexed and cited the same way.

Does Nexus need a cloud LLM to work?

No. Each of its four model roles — embed, distill, rerank, synthesize — is configured separately, and all four can run locally through Ollama or LM Studio. You can also mix them: a local embedder for private data paired with a stronger hosted model for the final answer.

How is this different from just adding a vector database in front of my docs?

A single embedding search misses exact strings — error codes, flag names, IDs — that a keyword search catches, and it ranks short filler messages competitively with long, correct answers. Nexus runs up to seven retrievers per query, fuses them with reciprocal rank fusion, and reranks the result against the actual question before answering, so no single technique’s blind spot decides the answer.

Is it safe to point at real company data?

Not without a security review first, and the project says so. Source credentials are Fernet-encrypted at rest, access follows OIDC and per-project roles, and every retrieved answer is scoped to sources the asker can already see — but that’s a starting point for a review, not a substitute for one. Start with public or low-sensitivity sources, and get Security and Legal to sign off before connecting anything else.

What does it cost to run?

Every LLM call is logged, with per-user daily token budgets and an instance-wide dollar ceiling that warns at 80% and stops at 100%. Running everything locally on one machine costs hardware time only; the bill only shows up if you choose a hosted model for one of the four capabilities.