Two datasets you can build on today, free.
The open Library and the Domain Index share the schema of everything we sell. Prototype on them, cite them, ship with them. When you need more text, license a shelf.
The open Library
Ten thousand complete, distilled documents across all eight shelves, each with its source, Bear Rank, word and token counts, reading level and cleanliness score. Every document is a page you can read, a Markdown file you can fetch and a line of JSONL you can load.
The Domain Index
One million domains with Bear Rank, its four sub-scores, tier, category, family-safe flag and the raw public signals behind them. It is the score attached to every document in the corpus, published so you can filter or weight by source quality yourself.
Attribution: "Directory Bear" with a link. No keys, no sign-up, no rate limits beyond the CDN's.
Your purchases, one command.
Every purchase comes with a customer ID. Install the dirbear client, enter the ID, and browse and download what you own — one file or everything — with resumable transfers and a checksum on every file. Files arrive as JSONL.zst; the client converts them to Parquet on your machine with one keystroke. Your ID keeps working for twelve months, and any new build we release in that time simply appears in your list.
pipx install dirbear dirbear
Scripting a server? This downloads and converts everything without the menu.
dirbear --customer DB-… --all --parquet
{ ≡ } Dir Bear · secure downloads ────────────────────────────────────────────────────────── › Customer ID DB-7F3A-9C21-XK4Q Product Files Size Build Finance & Business 42 61.2 GB 2026.09 Mixed bundles 18 38.0 GB 2026.09 Customer DB-7F3A-9C21-XK4Q · Acme AI · access until 2027-09-23 · 60 files, 99.2 GB › What would you like to do? ❯ Browse files and download one Download everything Download everything and convert to Parquet Convert downloaded files to Parquet (~/DirBear) Exit finance-2025-26-003.jsonl.zst ━━━━━━━━━━━━━━━━━━━━━━━━━━ 1.4/1.9 GB 84 MB/s 0:00:06 wrote finance-2025-26-002.parquet (61,240 rows, 1.7 GB)
Try a slice before you buy a shelf.
Ready on the CDN today, same schema as everything else, delivered the moment you check out.
Ready-made training corpora, one shelf at a time.
Each shelf is a complete, cleaned corpus: hundreds of gigabytes of full-length documents from domains that passed every gate, with Bear Rank and an exact token count on every row. Buy one shelf for a domain-specific model, or several and blend your own mix. Download links arrive within minutes.
Prices in USD, one-time, with two months of download access. Sizes are compressed; expect roughly 3× on disk uncompressed. Commercial training and fine-tuning included; redistribution of the raw text is not.
The entire distilled corpus. One fee, a year of access.
All shelves, every document, full metadata, as JSONL with Parquet conversion built into the client. Every build published during your year is included, each with a manifest and drop-rate report.
Every tier, one schema.
| Included | Open Library | Domain Index | Shelf corpus | Whole Library |
|---|---|---|---|---|
| Full distilled document text | ✓ 10k | — | ✓ | ✓ |
| Formats | HTML · MD · JSONL | Parquet · CSV | JSONL.zst · Parquet | JSONL.zst · Parquet |
| Crawl dump + date on every row | ✓ | — | ✓ | ✓ |
| Bear Rank + sub-scores per source | ✓ | ✓ | ✓ | ✓ |
| Tokens, words, reading level, cleanliness | ✓ | — | ✓ | ✓ |
| Raw domain signals (traffic, links, age) | — | ✓ | — | ✓ |
| Build manifest + drop-rate report | ✓ | ✓ | ✓ | ✓ |
| Rebuild cadence | — | — | quarterly | quarterly |
| Licence | CC BY 4.0 | CC BY 4.0 | commercial | commercial + internal redistribution |
The practical stuff.
Where does the text come from?
A variety of public web crawls and other data sources. We don't resell a crawl; we distil it. The value is in the gates: which domains are allowed in, what is stripped, what is deduplicated, and the counts attached to what is left.
Why full documents instead of snippets?
Because models learn from the whole argument, not the first paragraph. Every document in the corpus is at least 300 words of prose and most are far longer. Nothing is truncated.
How big is it, really?
A build starts from roughly 50 TB of raw crawl and ends at about 2.8 TB compressed, ~41 million documents. The size is printed on every product card and repeated in the manifest inside each download.
Can I train commercial models on it?
Yes for shelf corpora and the Whole Library. The Library and Domain Index are CC BY 4.0, which also allows commercial use with attribution. Source pages keep their own copyright; we supply extracted, transformed data and metadata.
What formats?
Every file is delivered as zstd-compressed JSONL, one document per line — readable by every tool and streamable at any size. The download client converts any file to Parquet locally, so you get the dataframe format without a second download. Row counts, byte sizes and total tokens are in the manifest.
How do updates work?
Updates come with the Whole Library: twelve months of access, and every new build we release in that window appears in your client automatically — no re-purchase, no re-download of what you already have. Individual shelves are a one-time download of the current build, with two months to fetch it. Each build is date-stamped so you know which crawl and which signal snapshot it came from.
Clean text, at the scale you're training at.
Whether you need one shelf for a domain model or the whole corpus with a year of fresh builds, we'll get you set up with a customer ID the same day. Invoices, licence questions, or just want to talk it through — write to us. We usually reply within two business days.
hello@dirbear.com