License operational data that never appeared on the open web
TGDC licenses rights-cleared operational data from real European companies to AI labs: coherent document archives, structured extracts, and held-out evaluation sets. Every corpus arrives with written chain of title and a per-batch anonymization report.
- Formats
- PDF · EML · XLSX · DOCX for archives, JSONL · Parquet for extracts
- Provenance
- Written mandate at the source, documented chain of custody
- Per batch
- Anonymization report and checksummed manifest, every delivery
What you can license
- Archives. Coherent document estates: correspondence, contracts, invoices and ledgers spanning years of real operations. Delivered as PDF, EML, XLSX and DOCX.
- Structured extracts. Entity-resolved, anonymized extractions delivered against your schema, as JSONL or Parquet.
- Evaluation sets. Held-out, never-published document tasks for benchmarking extraction, reasoning and agent workflows.
- RL environments. Task-based training environments with programmatic verifiers, built from the same cleared sources. See RL environments.
- Bespoke corpora. Custom professional datasets assembled on request across healthcare, legal and finance.
Provenance is the product
Every corpus enters through an authorized channel: insolvency estates, which administrators release only under a written mandate, workflow recordings made with consent, while the work is happening, and personal documents contributed by their owners with on-device anonymization. Nothing is scraped, and no corpus ships without a rights review covering ownership, third-party material and statutory limits.
The result is data with a paper trail: who owned it, who authorized it, what was removed, and what license terms attach. For labs facing training-data documentation obligations under the EU AI Act, that paper trail is the difference between an asset and a liability.
Anonymization, verified per batch
Named entities, personal data and commercial identifiers are removed or pseudonymized before delivery. Pseudonymization is deterministic, so cross-document references stay coherent: the same counterparty maps to the same pseudonym across an entire estate, without exposing who that counterparty was. Each batch ships with a verification report documenting what was detected and replaced.
Why operational data
Web text is exhausted and contaminated: whatever is publicly crawlable is already inside every frontier model. The operational record of a company, its correspondence, contracts, invoices and disputes over years, never appeared online. It carries the longitudinal structure, messiness and domain depth that scraped data cannot provide.
How delivery works
Every corpus follows the same four steps: acquired only under written authorization, cleared through a rights review, anonymized and verified per batch, and delivered as staged corpora with manifests, checksums and license terms. License terms are set per agreement, including exclusive licensing at corpus level.
What we send a seller
Two one-page briefs go to companies considering whether to license their records. They state what we ask for, how files reach us, what is replaced before delivery, and how a seller is paid. If you do not hold records yourself but can introduce us to a company that does, or bring a partnership our way, there is a referral fee in it for you: terms are agreed case by case before anything moves.
Scope a corpus. Tell us the domain and format you need: hello@thegeneraldata.com. Common questions are answered in the FAQ.