How we know what we know
The model, the method and its limits.
Everything we publish — every tender page, every signal, every figure in the Daily Tender — is produced by a data model and a culture of verification. This page is both, in full.
1 · The model: a graph of who buys what from whom
The corpus is not a pile of notices but a connected graph built from them:
- 19.3 million notices — the events: who required what and when, under which procedure, above or below the EU threshold.
- 13.7 million contract awards — the outcomes: winner, price, date, linked back to their notice.
- 1.7 million companies — the winners as legal entities, 1.2 million of them bound to official registry numbers (NIF, SIREN, CUI, EIK, Y-tunnus …), so that namesakes never merge and subsidiaries never blur.
- Normalised contracting authorities × CPV categories × years — each authority's procurement rhythm, machine-readable.
Most procurement data stops at the first point. The intelligence sits in the connections: the same records, linked, answer questions they cannot answer on their own.
2 · The method: gates instead of gut feeling
The pipeline was rebuilt around hard verification gates after we caught it, in its first months, silently losing fields at scale. The rules since then:
- Never guess a field. Every source integration starts from the portal's actual responses; every field is either mapped or explicitly skipped with a stated reason.
- One record in full before a million run. Dry runs compare source against storage, field by field, before every bulk run.
- Unknown stays unknown. The suitability check answers met / not met / unknown — missing information is never turned into a verdict.
- Nothing goes out on a single measurement. Machine-detected signals are cross-checked by an independent second method (a different comparison statistic, per-country normalisation, a recount after deduplication). What fails is discarded — and the discarding is logged.
- We audit ourselves in public. Our own mistakes become articles; the pipeline's pulse is a public page.
3 · What the model can do today
- Find — every open tender from 27 countries, searchable in English, including the 25,615 below-threshold procedures no aggregator lists.
- Qualify — contracting authorities' selection criteria, structured for 48,470 procedures; traceable suitability verdicts against a company profile.
- Contextualise — for every tender: the authority's history, the field of incumbents, what was paid before.
- Anticipate — 54–62% of authority × category pairs repeat year after year (measured); the cycle layer turns that into a forward calendar.
- Detect — daily anomalies with historical comparison cases, independently cross-checked before publication.
- Investigate — just-below-threshold ratios, event correlations (floods, wars, budget cycles), concentration patterns — the class of findings behind the Daily Tender.
4 · What we deliberately do not claim
- No win probabilities in the product. We built an experimental ML model; its honest metric is far below anything we would sell. Until a model beats traceable context by a publishable margin, it stays in the lab. (Vendors quoting impressive hit rates rarely publish their evaluation design. Ask them for it.)
- No causality from correlation. Our event articles claim temporal association, not cause — each one's methods box states exactly what was measured.
- No silent coverage promises. Sources publish differently: some only retrospectively, some lack categories. Where coverage is thin, our figures say so — they are lower bounds, not upper bounds.
- Translation is an index, not a legal text. Original titles are preserved throughout; contracts are won on the source documents.
The corpus behind this method: the data page. The method applied daily: The Daily Tender.
SOURCE · all figures on this page measured on our own corpus, last 19 August at 19:51