RETRIEVAL · SEE VOXELL FORGE IN ACTION ON PUBLIC DOCUMENTS

MEASURED ON 7,817 REAL DOCUMENTS AND 980,885 PASSAGES: THE FIRST RESULT ANSWERS THE QUESTION 83% OF THE TIME, AND ONE OF THE TOP THREE DOES 90% OF THE TIME.

SEC filings, arXiv technical papers, NASA technical reports and USPTO patents. 800 questions, scored on the product's own retrieval receipt. 80 of them are published here with their results, including the ones it got wrong.

How to read this. A model wrote the questions and their reference answers from the documents, and a model judged the top three results against them. That is easier than a test set written by people. These figures describe this system on these four corpora. They are not a comparison with any other vendor, and not a promise about your corpus. How it is measured.

nDCG@10 runs from 0 to 1 and rewards putting the right passage near the top of ten results. Each figure is the mean over about 200 questions, with the range a bootstrap puts it in 95% of the time. "Answers the question" is a model judge's reading of the returned passage against the question's reference answer. Across the four corpora the right document is in the top ten 95% of the time. How it is measured.

METHOD · IN PLAIN WORDS

HOW IT IS MEASURED.

Each corpus has 200 questions. Each one names its subject, the company or paper or patent it is about, and each has one passage labelled as the answer. The product retrieves ten passages per question and the receipt scores where the labelled passage landed. A question whose document could not be read is left out, and the receipt says how many were scored.

The headline figures use a second check on the same questions. A model judge reads each of the first three passages against the question's reference answer and says whether the passage states the facts that answer relies on, whether or not it is the labelled passage. "The first result answers the question" means the judge said yes to result one. A question with no stored reference answer is not judged. The headline is the mean of the four corpora's rates, weighted by their question counts, measured 2026-10-05.

Read these as what this system did on these questions. The questions are model-written, which makes them easier than a test set written by people. They are the right tool for telling one configuration from another on the same questions. They are not a promise about your corpus, which is why the offer below starts by measuring yours.

SEC FILINGS · SAME 200 QUESTIONS AT EVERY STEP

WHAT MOVED THE NUMBER.

01 · NDCG@10 0.467

Cross-encoder reranker

A cross-encoder rereads the top candidates against the question and reorders them.

02 · NDCG@10 0.622

Document identity in every passage

Each passage carries a one-line identity of the filing it came from: company, form, period.

03 · NDCG@10 0.843

Routing to the documents a question names

When a question names a company or filing, retrieval searches those documents first.

04 · NDCG@10 0.841

Voxell's own reranker

A reranker trained in-house replaced the larger general-purpose one. It rereads the same candidates at about a fifth of the compute per question.

SEC filings are the hard case: thousands of documents that share the same forms and the same boilerplate, where the right answer depends on whose filing it is. Read the SEC receipt →

CORPUS HEALTH · BEFORE ANYONE ASKS A QUESTION

DOCUMENT QUALITY IS GRADED.

Retrieval can only return what the corpus holds, so each document is scored on its own terms, with no question and no answer key. See how corpus scoring works →

RETRIEVAL EXPERTS · THE SAME YARDSTICK, ON YOUR CORPUS

THE MODEL IS NOT YOUR BOTTLENECK.

Voxell's Ingot-8B-R3 ranks #1 for English on the public MTEB leaderboard (English v2), with a 75.98 mean task score across 41 tasks. We are the people with the most to gain from telling you the model is the answer, and the steps above say otherwise: none of them changed the model.

Recall quality lives in the whole system: how documents are parsed, how structure survives ingestion, how chunks form, how queries are read, how results are ranked, and whether any of it is honestly measured. That is the work we take on.

TRANSMISSION · FROM THE FOUNDER

THE PERSON ON THIS PAGE IS THE PERSON IN YOUR REPO.

I am Jonathan Corners. Thirty years of enterprise systems, then Ingot. There is no sales team here and no delivery bench behind a curtain. Send a sample and I read it. Engage and I do the work.

I will find where your retrieval loses answers, show you what each miss costs, and prove the fix on the yardstick that found the problem.

Best fit: you already have retrieval in production and can name where it fails. Still building? Start with Answers instead.

JONATHAN CORNERS · FOUNDER, VOXELL

SCOPE · WHERE ANSWERS GET LOST

WHERE ANSWERS GET LOST. AND HOW WE GET THEM BACK.

01 · RAG PATTERNS, AT SCALE

PIPELINES THAT SURVIVE THEIR OWN CORPUS.

Naive top-k over flat chunks works in the demo and dies in production. We design and tune the full pattern range: standard vectorization done right, hierarchical enrichment, RAPTOR-style recursive abstraction, hybrid dense plus lexical retrieval, and reranking that earns its latency.

RAPTORHIERARCHICAL ENRICHMENTHYBRID RETRIEVALRERANKING

02 · TABLES AND TIME

YOUR ANSWERS LIVE IN TABLES. VECTORS WERE BUILT FOR PROSE.

Metrics, ledgers, and time series do not embed like paragraphs, and pretending they do is why a question about Q3 churn returns a paragraph near the number instead of the number. We make tabular and temporal data first-class citizens of retrieval.

TIME-SERIES RETRIEVALTABULAR INTEGRATIONSTRUCTURED GROUNDING

03 · CODE-AWARE INGEST

CODE IS NOT TEXT.

Splitting a function in half produces context that compiles nowhere. We ingest code the way compilers read it: AST-aware chunking that keeps functions, signatures, and call structure intact, so retrieval returns context that runs.

AST CHUNKINGSTRUCTURE-AWARE INGEST

04 · DOCUMENT STRUCTURE

DOCUMENTS HAVE BONES. MOST INGESTION CRUSHES THEM.

Most pipelines write their enrichment to a metadata column that retrieval never reads. Ours goes into the vector: the embedding is computed from the chunk plus its summary and named entities, while the stored text stays the raw chunk — so a query for "magic items" finds the gear line that never uses the phrase. Hypothetical questions are embedded as their own retrievable units, and a reduce pass builds document-level context above the chunks.

ENRICHMENT IN THE VECTORHYDE UNITSDOCUMENT-LEVEL SYNTHESIS

05 · MEASUREMENT AND JUDGING

EVALS THAT WOULD SURVIVE PEER REVIEW.

Recall@10 and nDCG@10 on golden sets. LLM judges calibrated against position, verbosity, and self-preference bias. Locked baselines, so every improvement is a delta, not an anecdote. This is the discipline the rest of the work stands on. Unmeasured work is redecorating.

RECALL@10NDCG@10CALIBRATED JUDGINGGOLDEN SETSLOCKED BASELINES

THE MOTION · PROOF BEFORE PROCUREMENT

YOU RISK A SAMPLE. WE RETURN A NUMBER.

01 · SEND A SAMPLE

A SLICE OF YOUR CORPUS. YOUR HARDEST QUERIES.

A sanitized slice of your corpus and 10 to 50 real queries, the ones your system gets wrong. NDA first if you want one.

02 · GET THE TEARDOWN

A SCORED BASELINE, IN DAYS.

Your Recall@10 and nDCG@10 on a receipt like the ones above, the questions your system answered wrong, and what each miss costs in your terms.

FREE · NO COMMITMENT

03 · THE FIX

FIXED SCOPE. NAMED PATH. SAME YARDSTICK.

An engagement with the remediation path named up front. Every change is measured against your locked baseline. You see the same numbers we do.

PRICED PER ENGAGEMENT · SCOPE AND PRICE CONFIRMED AFTER THE TEARDOWN

04 · KEEP IT FIXED

RAG-OPS, THE STANDING SERVICE.

Continuous evaluation and regression watch across every model, prompt, and corpus change. Retrieval quality decays silently. We watch the gauges.

PRICED PER ENGAGEMENT · ASK FOR A QUOTE

SPECIALIZED · BY DESIGN

WE DO ONE THING.

No transformation practice. No app-dev bench. No fourteen service lines. Retrieval and recall, all day, every day. Specialization is why the receipts exist.

PROOF · NOT PROMISES

SEND US A SAMPLE.

Share a link to a representative slice and the questions it should answer. Large corpus? Tell us its size and we will arrange a transfer. You get numbers and the answers you are missing. Then you decide.

NO INTRO CALL REQUIRED · NDA ON REQUEST · SANITIZED SLICES WELCOME