RETRIEVAL · SEE VOXELL FORGE IN ACTION ON PUBLIC DOCUMENTS
MEASURED ON 7,817 REAL DOCUMENTS AND 980,885 PASSAGES: THE FIRST RESULT ANSWERS THE QUESTION 83% OF THE TIME, AND ONE OF THE TOP THREE DOES 90% OF THE TIME.
SEC filings, arXiv technical papers, NASA technical reports and USPTO patents. 800 questions, scored on the product's own retrieval receipt. 80 of them are published here with their results, including the ones it got wrong.
How to read this. A model wrote the questions and their reference answers from the documents, and a model judged the top three results against them. That is easier than a test set written by people. These figures describe this system on these four corpora. They are not a comparison with any other vendor, and not a promise about your corpus. How it is measured.
nDCG@10 runs from 0 to 1 and rewards putting the right passage near the top of ten results. Each figure is the mean over about 200 questions, with the range a bootstrap puts it in 95% of the time. "Answers the question" is a model judge's reading of the returned passage against the question's reference answer. Across the four corpora the right document is in the top ten 95% of the time. How it is measured.
METHOD · IN PLAIN WORDS
HOW IT IS MEASURED.
Each corpus has 200 questions. Each one names its subject, the company or paper or patent it is about, and each has one passage labelled as the answer. The product retrieves ten passages per question and the receipt scores where the labelled passage landed. A question whose document could not be read is left out, and the receipt says how many were scored.
The headline figures use a second check on the same questions. A model judge reads each of the first three passages against the question's reference answer and says whether the passage states the facts that answer relies on, whether or not it is the labelled passage. "The first result answers the question" means the judge said yes to result one. A question with no stored reference answer is not judged. The headline is the mean of the four corpora's rates, weighted by their question counts, measured 2026-10-05.
Read these as what this system did on these questions. The questions are model-written, which makes them easier than a test set written by people. They are the right tool for telling one configuration from another on the same questions. They are not a promise about your corpus, which is why the offer below starts by measuring yours.
SEC FILINGS · SAME 200 QUESTIONS AT EVERY STEP
WHAT MOVED THE NUMBER.
01 · NDCG@10 0.467
Cross-encoder reranker
A cross-encoder rereads the top candidates against the question and reorders them.
02 · NDCG@10 0.622
Document identity in every passage
Each passage carries a one-line identity of the filing it came from: company, form, period.
03 · NDCG@10 0.843
Routing to the documents a question names
When a question names a company or filing, retrieval searches those documents first.
04 · NDCG@10 0.841
Voxell's own reranker
A reranker trained in-house replaced the larger general-purpose one. It rereads the same candidates at about a fifth of the compute per question.
SEC filings are the hard case: thousands of documents that share the same forms and the same boilerplate, where the right answer depends on whose filing it is. Read the SEC receipt →
CORPUS HEALTH · BEFORE ANYONE ASKS A QUESTION
DOCUMENT QUALITY IS GRADED.
Retrieval can only return what the corpus holds, so each document is scored on its own terms, with no question and no answer key. See how corpus scoring works →
RETRIEVAL EXPERTS · THE SAME YARDSTICK, ON YOUR CORPUS
THE MODEL IS NOT YOUR BOTTLENECK.
Voxell's Ingot-8B-R3 ranks #1 for English on the public MTEB leaderboard (English v2), with a 75.98 mean task score across 41 tasks. We are the people with the most to gain from telling you the model is the answer, and the steps above say otherwise: none of them changed the model.
Recall quality lives in the whole system: how documents are parsed, how structure survives ingestion, how chunks form, how queries are read, how results are ranked, and whether any of it is honestly measured. That is the work we take on.
TRANSMISSION · FROM THE FOUNDER
THE PERSON ON THIS PAGE IS THE PERSON IN YOUR REPO.
I am Jonathan Corners. Thirty years of enterprise systems, then Ingot. There is no sales team here and no delivery bench behind a curtain. Send a sample and I read it. Engage and I do the work.
I will find where your retrieval loses answers, show you what each miss costs, and prove the fix on the yardstick that found the problem.
Best fit: you already have retrieval in production and can name where it fails. Still building? Start with Answers instead.
SCOPE · WHERE ANSWERS GET LOST
WHERE ANSWERS GET LOST. AND HOW WE GET THEM BACK.
01 · RAG PATTERNS, AT SCALE
PIPELINES THAT SURVIVE THEIR OWN CORPUS.
Naive top-k over flat chunks works in the demo and dies in production. We design and tune the full pattern range: standard vectorization done right, hierarchical enrichment, RAPTOR-style recursive abstraction, hybrid dense plus lexical retrieval, and reranking that earns its latency.
02 · TABLES AND TIME
YOUR ANSWERS LIVE IN TABLES. VECTORS WERE BUILT FOR PROSE.
Metrics, ledgers, and time series do not embed like paragraphs, and pretending they do is why a question about Q3 churn returns a paragraph near the number instead of the number. We make tabular and temporal data first-class citizens of retrieval.
03 · CODE-AWARE INGEST
CODE IS NOT TEXT.
Splitting a function in half produces context that compiles nowhere. We ingest code the way compilers read it: AST-aware chunking that keeps functions, signatures, and call structure intact, so retrieval returns context that runs.
04 · DOCUMENT STRUCTURE
DOCUMENTS HAVE BONES. MOST INGESTION CRUSHES THEM.
Most pipelines write their enrichment to a metadata column that retrieval never reads. Ours goes into the vector: the embedding is computed from the chunk plus its summary and named entities, while the stored text stays the raw chunk — so a query for "magic items" finds the gear line that never uses the phrase. Hypothetical questions are embedded as their own retrievable units, and a reduce pass builds document-level context above the chunks.
05 · MEASUREMENT AND JUDGING
EVALS THAT WOULD SURVIVE PEER REVIEW.
Recall@10 and nDCG@10 on golden sets. LLM judges calibrated against position, verbosity, and self-preference bias. Locked baselines, so every improvement is a delta, not an anecdote. This is the discipline the rest of the work stands on. Unmeasured work is redecorating.
THE MOTION · PROOF BEFORE PROCUREMENT
YOU RISK A SAMPLE. WE RETURN A NUMBER.
01 · SEND A SAMPLE
A SLICE OF YOUR CORPUS. YOUR HARDEST QUERIES.
A sanitized slice of your corpus and 10 to 50 real queries, the ones your system gets wrong. NDA first if you want one.
02 · GET THE TEARDOWN
A SCORED BASELINE, IN DAYS.
Your Recall@10 and nDCG@10 on a receipt like the ones above, the questions your system answered wrong, and what each miss costs in your terms.
FREE · NO COMMITMENT
03 · THE FIX
FIXED SCOPE. NAMED PATH. SAME YARDSTICK.
An engagement with the remediation path named up front. Every change is measured against your locked baseline. You see the same numbers we do.
PRICED PER ENGAGEMENT · SCOPE AND PRICE CONFIRMED AFTER THE TEARDOWN
04 · KEEP IT FIXED
RAG-OPS, THE STANDING SERVICE.
Continuous evaluation and regression watch across every model, prompt, and corpus change. Retrieval quality decays silently. We watch the gauges.
PRICED PER ENGAGEMENT · ASK FOR A QUOTE
SPECIALIZED · BY DESIGN
WE DO ONE THING.
No transformation practice. No app-dev bench. No fourteen service lines. Retrieval and recall, all day, every day. Specialization is why the receipts exist.
PROOF · NOT PROMISES
SEND US A SAMPLE.
Share a link to a representative slice and the questions it should answer. Large corpus? Tell us its size and we will arrange a transfer. You get numbers and the answers you are missing. Then you decide.
NO INTRO CALL REQUIRED · NDA ON REQUEST · SANITIZED SLICES WELCOME