Open data

Two datasets you can build on today, free.

The open Library and the Domain Index share the schema of everything we sell. Prototype on them, cite them, ship with them. When you need more text, license a shelf.

Free full text

The open Library

Ten thousand complete, distilled documents across all eight shelves, each with its source, Bear Rank, word and token counts, reading level and cleanliness score. Every document is a page you can read, a Markdown file you can fetch and a line of JSONL you can load.

Documents10,000Words18.4MShelves8Size~180 MBFormatsHTML · MD · JSONLLicenceCC BY 4.0
Free domain scores

The Domain Index

One million domains with Bear Rank, its four sub-scores, tier, category, family-safe flag and the raw public signals behind them. It is the score attached to every document in the corpus, published so you can filter or weight by source quality yourself.

Rows1,000,000Columns14Domains scoredall TLDsSize~480 MBFormatsParquet · CSVLicenceCC BY 4.0

Attribution: "Directory Bear" with a link. No keys, no sign-up, no rate limits beyond the CDN's.

Delivery

Your purchases, one command.

Every purchase comes with a customer ID. Install the dirbear client, enter the ID, and browse and download what you own — one file or everything — with resumable transfers and a checksum on every file. Files arrive as JSONL.zst; the client converts them to Parquet on your machine with one keystroke. Your ID keeps working for twelve months, and any new build we release in that time simply appears in your list.

pipx install dirbear
dirbear

Scripting a server? This downloads and converts everything without the menu.

dirbear --customer DB-… --all --parquet
dirbear
  { ≡ }  Dir Bear   ·   secure downloads
  ──────────────────────────────────────────────────────────

 Customer ID  DB-7F3A-9C21-XK4Q

  Product                 Files      Size    Build
  Finance & Business         42   61.2 GB    2026.09
  Mixed bundles              18   38.0 GB    2026.09
  Customer DB-7F3A-9C21-XK4Q · Acme AI · access until 2027-09-23 · 60 files, 99.2 GB

 What would you like to do?
  ❯ Browse files and download one
    Download everything
    Download everything and convert to Parquet
    Convert downloaded files to Parquet  (~/DirBear)
    Exit

  finance-2025-26-003.jsonl.zst  ━━━━━━━━━━━━━━━━━━━━━━━━━━  1.4/1.9 GB  84 MB/s  0:00:06
  wrote   finance-2025-26-002.parquet  (61,240 rows, 1.7 GB)
Starter sets

Try a slice before you buy a shelf.

Ready on the CDN today, same schema as everything else, delivered the moment you check out.

Licensed corpora · by the shelf

Ready-made training corpora, one shelf at a time.

Each shelf is a complete, cleaned corpus: hundreds of gigabytes of full-length documents from domains that passed every gate, with Bear Rank and an exact token count on every row. Buy one shelf for a domain-specific model, or several and blend your own mix. Download links arrive within minutes.

Prices in USD, one-time, with two months of download access. Sizes are compressed; expect roughly 3× on disk uncompressed. Commercial training and fine-tuning included; redistribution of the raw text is not.

Whole Library

The entire distilled corpus. One fee, a year of access.

All shelves, every document, full metadata, as JSONL with Parquet conversion built into the client. Every build published during your year is included, each with a manifest and drop-rate report.

Whole Library · 12 monthsseveral terabytes compressed, all shelves and mixed bundles, every new build released in your year, email support
$8,000 one-time
Every new build, includedEach crawl we distil during your year appears in your client automatically — no re-purchase, no re-download of what you already have
Commercial licenceTrain, fine-tune, evaluate and ship models on it; the only thing you may not do is resell the raw text
Talk to us
What's in the files

Every tier, one schema.

IncludedOpen LibraryDomain IndexShelf corpusWhole Library
Full distilled document text✓ 10k
FormatsHTML · MD · JSONLParquet · CSVJSONL.zst · ParquetJSONL.zst · Parquet
Crawl dump + date on every row
Bear Rank + sub-scores per source
Tokens, words, reading level, cleanliness
Raw domain signals (traffic, links, age)
Build manifest + drop-rate report
Rebuild cadencequarterlyquarterly
LicenceCC BY 4.0CC BY 4.0commercialcommercial + internal redistribution
Questions

The practical stuff.

Where does the text come from?

A variety of public web crawls and other data sources. We don't resell a crawl; we distil it. The value is in the gates: which domains are allowed in, what is stripped, what is deduplicated, and the counts attached to what is left.

Why full documents instead of snippets?

Because models learn from the whole argument, not the first paragraph. Every document in the corpus is at least 300 words of prose and most are far longer. Nothing is truncated.

How big is it, really?

A build starts from roughly 50 TB of raw crawl and ends at about 2.8 TB compressed, ~41 million documents. The size is printed on every product card and repeated in the manifest inside each download.

Can I train commercial models on it?

Yes for shelf corpora and the Whole Library. The Library and Domain Index are CC BY 4.0, which also allows commercial use with attribution. Source pages keep their own copyright; we supply extracted, transformed data and metadata.

What formats?

Every file is delivered as zstd-compressed JSONL, one document per line — readable by every tool and streamable at any size. The download client converts any file to Parquet locally, so you get the dataframe format without a second download. Row counts, byte sizes and total tokens are in the manifest.

How do updates work?

Updates come with the Whole Library: twelve months of access, and every new build we release in that window appears in your client automatically — no re-purchase, no re-download of what you already have. Individual shelves are a one-time download of the current build, with two months to fetch it. Each build is date-stamped so you know which crawl and which signal snapshot it came from.

Contact

Clean text, at the scale you're training at.

Whether you need one shelf for a domain model or the whole corpus with a year of fresh builds, we'll get you set up with a customer ID the same day. Invoices, licence questions, or just want to talk it through — write to us. We usually reply within two business days.

hello@dirbear.com