"pricing":
Pay per dataset. Verify before you pay at all.
Paid plans are per dataset, monthly, with no long-term lock-in. Sample access opens soon — until then, talk to us and we will walk you through the provenance chain on a real release.
demo
$0
sample access opens soon
The evaluation sample is not open yet. Talk to us in the meantime and we will walk you through a real release.
- Evaluation sample, opening soon
- Full provenance chain on every record
- Merkle root over every chunk hash
- The same JSONL the paid plans ship
starter
$99/mo
per dataset / month
A production corpus, refreshed weekly.
- Weekly release updates
- JSONL + Merkle-rooted manifest
- Per-record provenance fields
- Email support
pro
$299/mo
per dataset / month
Daily releases with stable IDs — re-embed only what moved.
- Daily release updates
- Stable record and chunk IDs across releases
- RAG/embedding license with indemnification
- Priority support
enterprise
Custom
custom agreement
Every dataset, your niches, our pipeline.
- Every dataset we publish
- Freshness and support terms in the contract
- Custom niches and connectors
- Dedicated support
every plan includes
Full per-record provenance chain
Merkle root over every chunk hash
JSONL + per-file SHA-256
Verifiable with your own code
Need every niche, or one we don't have?
Enterprise covers every dataset we publish, custom niches and connectors, dedicated support, and freshness terms agreed in the contract. Billing goes through a contract, not a checkout.
"faq":
Questions buyers actually ask
Can I legally embed and retrieve over this data?
Yes — that is the point. Every dataset ships under an explicit license that permits RAG, embedding, and retrieval use. Pro and Enterprise licenses include indemnification. No scraping gray zones: we collect from primary sources, honor robots.txt and opt-outs, and screen personal data out of every record before it is hashed.
How do I verify provenance?
Every record carries its source URL, fetch timestamp, and SHA-256 of the raw response. Every release ships a manifest with a SHA-256 for every file and a Merkle root over the ordered chunk hashes. Recompute the file hashes, rebuild the root from chunks.jsonl, compare — a few lines of your own code, none of ours. We ran exactly that against a full-history release with a separate implementation, and the roots matched.
How often are datasets updated?
Cadence is per dataset and per plan: Starter ships weekly releases, Pro ships daily. Each release is versioned, and because record and chunk IDs are stable across releases you can compare any two of them yourself and see exactly what moved.
What formats do you ship?
JSONL for records and retrieval-ready chunks, plus a manifest.json per release carrying every file's SHA-256 and byte count, the record and chunk counts, quality metrics, and the Merkle root over the ordered chunk hashes.
Can I try before paying?
Not from a sample dataset today — sample access is closed while the corpus grows, and it opens soon. In the meantime, tell us which corpus you are evaluating: we will walk you through a real release — the manifest, the per-file SHA-256 and the Merkle root — and you can recompute the root with your own code before any money changes hands.
Do you build custom niches?
On the Enterprise plan we build custom niches and connectors against your source list, with freshness terms agreed in the contract. Tell us what your retrieval layer is missing via the contact form.