merkleset

$ sha256sum chunks.jsonl ✓ matches manifest

RAG-ready data.
Provenance you can verify.

Cryptographically verifiable, licensed-for-AI datasets from primary sources. Every record hashed, every release Merkle-rooted, every ID stable across releases — so your retrieval layer stands on data you can defend.

sample access opens soon · full provenance chain on every record · primary sources only

us-federal-procurement@2026.08.10 / manifest.json✓ verifiable
4b1d9e2cc7a01f8eab52630de9f42d7b8c3a5e91f04d7ba6d2c83a6f9f2c
merkle_root9f2c41ab…c135e41a✓ verifiable release

"why_merkleset":

Data you can put in front of your legal team

Verifiable provenance

Every record carries its source URL, fetch timestamp and SHA-256. Every release ships a manifest with a Merkle root over every chunk hash — recompute it from the files you downloaded and check any record yourself.

source_url · fetched_at · sha256

Licensed for RAG

An explicit RAG/embedding-use license instead of a gray area, with indemnification on Pro and Enterprise. Your legal review is one page, not forty.

embedding_use: permitted

Stable IDs across releases

record_id is a UUID v5 of the dataset plus the natural key; chunk_id is a content address. The same chunk in two releases carries the same ID and the same hash, so you can compare releases yourself and re-embed only what moved.

record_id · chunk_id · stable across releases

Procurement and regulatory, live across jurisdictions

US federal procurement ships today, alongside legislation, rulings and regulator guidance from the US, the EU, the UK and other national jurisdictions. Courts, SEC filings and e-commerce stay on the roadmap. All collected from primary sources — never resold vendor feeds.

procurement ✓ regulatory ✓ courts · sec · ecommerce

"how_it_works":

Pipeline, not promises

  1. 01 collect

    We collect from primary sources

    Official bulk files and public APIs — from SAM.gov and the Federal Register to the EU Publications Office and national legislation portals. Every raw response is stored as its own evidence record and hashed at fetch time.

  2. 02 normalize

    We normalize and chunk with stable IDs

    Cleaned, structured, chunked for retrieval. Chunk IDs are content addresses, so an unchanged chunk keeps its ID across releases.

  3. 03 release

    You pull versioned releases via API

    JSONL plus a manifest per release, carrying every file's SHA-256 and the Merkle root. Recompute the root before a single byte enters your index.

"catalog":

Procurement and regulatory live, more niches planned. One license.

"pricing":

Prices on the page, per dataset

full pricing

demo

$0

sample access opens soon

  • Evaluation sample, opening soon
  • Full provenance chain on every record
  • Merkle root over every chunk hash
  • The same JSONL the paid plans ship
Talk to us about a sample

starter

$99/mo

per dataset / month

  • Weekly release updates
  • JSONL + Merkle-rooted manifest
  • Per-record provenance fields
  • Email support
Contact us to subscribe
most popular

pro

$299/mo

per dataset / month

  • Daily release updates
  • Stable record and chunk IDs across releases
  • RAG/embedding license with indemnification
  • Priority support
Contact us to subscribe

snapshot

$599

per dataset, one-time

  • Latest full published release: records, chunks, Merkle-rooted manifest
  • Perpetual license for internal RAG and embedding use
  • Re-download the exact purchased version for 12 months
  • Option: add 12 months of updates for $990 total
Contact us to purchase

enterprise

Custom

custom agreement

  • Every dataset we publish
  • Freshness and support terms in the contract
  • Custom niches and connectors
  • Dedicated support
Contact us

"faq":

Questions buyers actually ask

Can I legally embed and retrieve over this data?

Yes — that is the point. Every dataset ships under an explicit license that permits RAG, embedding, and retrieval use. Pro and Enterprise licenses include indemnification. No scraping gray zones: we collect from primary sources, honor robots.txt and opt-outs, and screen personal data out of every record before it is hashed.

How do I verify provenance?

Every record carries its source URL, fetch timestamp, and SHA-256 of the raw response. Every release ships a manifest with a SHA-256 for every file and a Merkle root over the ordered chunk hashes. Recompute the file hashes, rebuild the root from chunks.jsonl, compare — a few lines of your own code, none of ours. We ran exactly that against a full-history release with a separate implementation, and the roots matched.

How often are datasets updated?

Cadence is per dataset and per plan: Starter ships weekly releases, Pro ships daily. Each release is versioned, and because record and chunk IDs are stable across releases you can compare any two of them yourself and see exactly what moved.

What formats do you ship?

JSONL for records and retrieval-ready chunks, plus a manifest.json per release carrying every file's SHA-256 and byte count, the record and chunk counts, quality metrics, and the Merkle root over the ordered chunk hashes.

Can I try before paying?

Not from a sample dataset today — sample access is closed while the corpus grows, and it opens soon. In the meantime, tell us which corpus you are evaluating: we will walk you through a real release — the manifest, the per-file SHA-256 and the Merkle root — and you can recompute the root with your own code before any money changes hands.

Do you build custom niches?

On the Enterprise plan we build custom niches and connectors against your source list, with freshness terms agreed in the contract. Tell us what your retrieval layer is missing via the contact form.