Seven subsystems, four flows, and which stores are precious.
Seven subsystems in one organisation's runtime, and one of ours apart. The Product Backend, the Backend for short, owns the product and the question ledger, is the only thing the apps call, and is always one organisation's. The Platform Backend at the top right is ours and is the only thing with authority across organisations: it mints them, holds licences and metering, signs its own operators in against its own small table, and is content-blind, receiving aggregates on one telemetry channel and never querying the runtime. Identity owns who you are and which groups you are in. Documents owns the organisation's documents as they are now, each with one identity of ours, decided there, in two forms, the raw file and a unified structural form derived from it by parsers that live inside Documents, and the sets they are organised into. The Vector Engine owns the searchable form derived from them, pieces into vectors and the nearest for a question, named for what it does today. Retrieval owns the grants that join a group to a set, and it is the one door through which a person reads: the Backend hands it the session, and Retrieval resolves who this is and what they may read, asks the Vector Engine or Documents inside that scope, and records the search as an event. That is enforced by credential, not convention; the person-read doors accept only Retrieval. Import owns the connections to the sources and the cursors into them; the adapters own each source's protocol.
Four flows, and every line belongs to one. The ask: an app to the Backend; the Backend to Retrieval with the session and the question; Retrieval to Identity for who this is and her groups, then to Vector Engine with the scope; evidence back to the Backend, which has the model outside write the answer, has Retrieval open the cited instruction within the same scope, and writes one ledger row. After the answer has gone, Retrieval asks the same question once more over the whole organisation and records on its own side whether it would have been answered, so the dashboard can tell knowledge missing from access preventing; that result never reaches the person and is shown to a manager only within the manager's own scope. The import: a source, through its adapter, to Import, into Documents, the cursor committed once Documents has stored the change and not before. The derive: Vector Engine asking Documents what changed since its mark and rebuilding on its own clock, so getting a document in and making it searchable never wait on each other. The setup: the Knowledge Manager's calls through the Backend to each owner, users and groups to Identity, grants to Retrieval, sets to Documents, connections to Import. One more flow runs outside the software and closes the loop: the manager reads the dashboard and fixes the file where it lives.
Where the data is: a store under every box that holds state, and each says which kind. Precious, backed up: the ledger, the users and groups, the grants, the connections, and above all Documents, the bytes, their identity and their sets, which nothing can regenerate. Derived, rebuildable: Vector Engine, pieces and vectors, a function of Documents and a version, dropped and rebuilt rather than backed up. The adapters hold nothing and have no store. Two shared capabilities are drawn where they are called: Email between the two backends, with a line from each because those two call it and nothing else does, templates and a queue and delivery, with Identity still the one that mints a sign-in token; and Logger beside the Platform Backend, written by every box and read by our operators through the Platform, under the rule that a log line never carries content, which is what makes that read safe and lets the same Logger ship into a customer's building. Both are the shape of services already running in production for another product of ours. There is no translation service: a template or an interface string is translated by a model call inside the subsystem that owns it. Postgres throughout, pgvector inside Vector Engine's. Every search is a statistic: Retrieval creates one event per query with the permitted outcome and the engine version, completes it afterwards with the organisation-wide outcome, the Backend's ledger row shares its id and adds the answer and the friction tap, the person's verdict travels back through the Backend onto the same event, and the dashboard reads four states from the two: answered, knowledge gap, access gap, quality problem. Above the apps, the two people the product is for: the Reader, who asks, opens and rates; the Knowledge Manager, who organises access and sources and fixes the knowledge when the system shows a gap. The surfaces beside each role are what that role can reach today, never what it may do.