For agentic RAG developers
FORGE
TURBO · FREE Your first vector in 60 seconds. Pipe a local corpus straight into GPU memory withforge-cli over gRPC and Protobuf — no HTTP hop.
YOUR TURN · SAME CALL, YOUR CODE
Run it in the playground with no login, or paste it after you grab a key. Turbo is free. Prefer a terminal? forge-cli embed --text "..." is a public download →
THE GAP
The gap, in practice
87ms vs 300ms. Same question, same corpus size. 213ms is the difference between an agent that chains ten calls and one that times out.
Deploy anywhere
Hosted at api.voxell.ai. Embedded in your VPC from the AWS Marketplace. Or on-prem as a Docker container behind your firewall. Same schema in all three.
Zero trust
mTLS client identity is available on every tier — Ed25519 client certificates in place of bearer tokens, no shared secrets to rotate. Every tier is encrypted in transit with TLS 1.3. Zero data retention after the response. Learn about VQS corpus quality scoring.
Numbers, with receipts: 87ms is P50 end to end on owned NVIDIA DGX hardware, no shared queue. 300ms is the observed market average across hosted embedding APIs. Turbo outscores OpenAI's best embedding model on English MTEB; the flagship holds rank #1 at 75.98 Mean (Task), independently verified on the public leaderboard. Methodology in the docs.