# Dir Bear > The web, distilled for AI. Dir Bear turns tens of terabytes of public web crawl into clean, > full-length, token-counted English documents, each scored with Bear Rank (a 0–100 domain > quality score) and placed on one of eight subject shelves. Ten thousand documents are open; > the corpus is sold by the shelf or as a whole. https://dirbear.com · X: @directorybear ## Pages - https://dirbear.com/ — overview, the six gates, the record schema - https://dirbear.com/library/ — the open Library: browse 10,000 documents by shelf - https://dirbear.com/datasets/ — free datasets, shelf corpora, Whole Library, download client ## Open data (CC BY 4.0, attribution "Dir Bear, https://dirbear.com") - https://dirbear.com/library/all.jsonl.gz — all 10,000 Library documents, one JSON object per line, full text - https://dirbear.com/library/{slug}.md — one document: YAML front-matter + text - https://dirbear.com/library/index.json — shelves, counts, chunk layout used by the Library page - https://dirbear.com/index/bear-rank-1m.csv.gz — the Domain Index, 1,000,000 domains (also Parquet) - https://dirbear.com/library/NOTICE.txt — licence and attribution - https://huggingface.co/datasets/directorybear/open-library-10k — the Open Library on Hugging Face (Parquet + all.jsonl.gz) - https://huggingface.co/datasets/directorybear/bear-rank-1m — the Domain Index on Hugging Face (Parquet + csv.gz) - https://huggingface.co/directorybear — Dir Bear on Hugging Face ## Record schema (identical in JSONL, Markdown front-matter and the paid corpus) - id: stable id, blake2b of the source URL - url, domain: the source page and its host (www stripped) - dump: which crawl the page came from (e.g. 2025-26); crawled: YYYY-MM-DD - bear_rank: 0–100 domain score; tier: grizzly ≥ 90 · kodiak 75–89 · polar 50–74 · cub < 50 - shelf: one of the shelves below; shelf_confidence: 0–1 (below 0.35 → sold in mixed bundles) - title; words; tokens (cl100k tokenizer); reading_level (Flesch–Kincaid grade) - cleanliness: 0–100, share of the final text that is prose; family_safe: boolean - build: corpus build id (e.g. 2026.09); text: the complete document ## Shelves (subjects) - science-engineering — Science & Engineering: research notes, lab guides, engineering explainers - software-docs — Software & Documentation: reference docs, tutorials, engineering blogs - health — Health & Medicine: patient information, clinical explainers, public health - finance — Finance & Business: markets, accounting, management, company explainers - education — Education & Reference: course material, encyclopedic entries, study guides - home-craft — Home, Food & Craft: cooking, gardening, repair, making things by hand - history-culture — History & Culture: essays, histories, arts and place writing - civic-law — Civic & Law: government guidance, legal explainers, policy - general — documents the classifier could not place with confidence (sold as mixed bundles) ## Domain categories (49, carried on the Domain Index as `category`; map to shelves as `shelf_prior`) Technology, Web Development, AI & Machine Learning, Cybersecurity, Cloud & Hosting, News & Media, Video & Streaming, Music & Audio, Gaming, Sports, Arts & Culture, Entertainment, Photography, Shopping, Automotive, Real Estate, Fashion & Beauty, Food & Drink, Business, Marketing & Advertising, Jobs & Careers, Legal, Finance, Cryptocurrency, Investing, Insurance, Education, Online Learning, Reference, Science, Books & Literature, Health, Fitness, Social Media, Communication, Forums & Communities, Search Engines, Government, Non-Profit, Religion, Travel, Home & Garden, Pets & Animals, Parenting & Family, Hobbies & Interests, Tools & Utilities, Weather, Maps & Directories, General ## Bear Rank - Inputs: Popularity 30 % (Tranco traffic rank, log scale) · Authority 30 % (Majestic referring subnets and IPs) · Longevity 20 % · Safety 20 % (TLS, security headers, abuse / adult / gambling lists) - Every source domain is scored; domains outside the ranked million receive a floor score (≈ 14, cub) ## The six gates (how a page earns its place) 1. Domain — scored with Bear Rank; floor BR 12 removes parked domains and link farms 2. Safety — domain on malware / phishing / adult / gambling / scam lists → dropped; content backstop 3. Extraction — residual boilerplate lines removed (cookie notices, nav, share bars, footers) 4. Substance — English, ≥ 250 words, ≤ 34 % list-like lines, prose-like; most of the web is short 5. Deduplication — exact copies collapsed (tracking-parameter and http/https twins, verbatim syndication) 6. Labelling — tokens, reading level, cleanliness, title, id, shelf with confidence ## Products - Open Library (10,000 documents) and Domain Index (1,000,000 domains): free, CC BY 4.0 - Shelf corpora: one-time licence per shelf, JSONL.zst with Parquet conversion in the client - Whole Library: one-time fee, twelve months of access, every new build in that year included - Delivery: the `dirbear` command-line client (pipx install dirbear) with a customer ID - Contact: hello@dirbear.com