Skip to content
diviteb

Answers with sources. Documents turned into data.

Retrieval over your docs, tickets, and code with cited answers — plus extraction pipelines that turn invoices, contracts, and forms into validated, structured records.

Two jobs, one pipeline

Six parts of a knowledge system worth shipping.

Retrieval answers questions. Extraction fills your systems. Both run on the same evals.

Hybrid retrieval with reranking

Vectors catch meaning. Keyword search catches exact terms, part numbers, and error codes. We run both, fuse the scores, rerank the top results, and tune the mix per corpus. Chunking follows the content: docs by heading, tickets by thread, code by function.

Permissions per user

Source permissions travel into the index. Results are filtered at query time for the person asking. If they cannot open a document, the system will not cite it for them.

SSO groupsRow-levelQuery-time

Citations on every answer

Each answer links the passages it drew from. Reviewers check the source in one click. No source, no answer — the system says it does not know.

InlineRankedAbstains

Evals for retrieval and faithfulness

Every change runs against a test set built from your real questions and documents. We score recall, faithfulness, and field accuracy against floors you agree to. A change below the floor does not merge.

Document extraction

Invoices, contracts, and forms in any layout become typed records. Every field is checked against your schema and your systems before it reaches the ERP or CRM.

Human review queues

Low-confidence fields go to a person with the source region highlighted. One keystroke accepts, one corrects. Corrections feed the next eval run.

The hard part

Catching hallucinations before users do.

A system that answers well most of the time still fails some of the time. The work is knowing which answers fail, why, and shipping a fix without breaking the rest.

  • Graded on correctness and groundedness. A right answer the sources do not support counts as a miss.
  • Failures are labelled: retrieval miss, missing context, generation drift, prompt injection.
  • Dashboards per failure type, so a fix targets the cause, not the symptom.
  • New failure cases from production join the test set every week.
Read the RAG guide

Extraction

Documents in, records out.

Invoices, contracts, and forms arrive in every layout. We extract to your schema, validate every field, and route doubtful documents to a person before anything is posted.

  • A typed schema per document type — fields, formats, and cross-field rules.
  • Checks against your systems: the PO exists, the totals add up, the vendor is known.
  • Confidence thresholds per field. Below the line, a person reviews it.
  • Every correction is logged with who made it and why.
  • Every answer

    Cites its sources, or says it does not know

  • Per user

    Permissions checked at retrieval time

  • Every change

    Retrieval and extraction evals run in CI

  • Your cloud

    Indexes, documents, and pipelines stay in your accounts

The review queue changed the job. Our clerks check the hard invoices instead of typing every one.
IllustrativeFinance Operations LeadLogistics company, 400 employees
See our work

Questions

What buyers ask us first.

Can people see answers from documents they can't open?
No. We carry your source permissions into the index and filter results per user at query time. A document someone cannot open is never cited for them.
Which documents can you extract from?
Anything with a stable meaning: invoices, purchase orders, contracts, claims, onboarding forms. Scans and handwriting work, with more of them routed to review. We test on a sample of your real documents before we commit to a scope.
Where does our data go?
Into your cloud accounts and the model providers you approve under your own agreements. When data cannot leave your environment, we run open-weight models inside it.
How long until it runs on real data?
We plan for a first retrieval or extraction pipeline on your data in four to six weeks. Evals run before any tuning starts, so progress is measured, not guessed.
How is it priced?
Each engagement is scoped and priced after a discovery call. Model, OCR, and hosting usage bills to your own accounts.

Ready when you are

Show us the documents your team retypes.

A 30-minute discovery call. Bring a sample. We'll tell you whether retrieval, extraction, or both fit.