How we know what we know
The model, the method, and its limits.
Everything we publish — every tender page, every signal, every Daily Tender figure — comes out of one data model and one verification culture. This page is the whole of it.
1 · The model: a graph of who buys what from whom
The corpus is not a pile of notices; it is a connected graph built from them:
- 18.6M notices — the events: who asked for what, when, under which procedure, above or below the EU threshold.
- 13.4M awards — the outcomes: winner, price, date, linked back to their notices.
- 1.6M companies — the winners as entities, 1.2M anchored to official registry identifiers (NIF, SIREN, CUI, EIK, Y-tunnus…) so namesakes never merge and subsidiaries never blur.
- Normalized buyers × CPV categories × years — every buyer’s purchasing rhythm, machine-readable.
Most procurement data stops at the first bullet. The intelligence lives in the joins: the same records, connected, answer questions the records alone cannot.
2 · The method: gates, not vibes
The pipeline was rebuilt around hard gates after we caught it silently losing fields at scale in its early months. The rules since:
- Never guess a field. Every source integration starts from the portal’s real responses; every field is mapped or explicitly skipped with a reason.
- One record end-to-end before a million. Dry-runs compare source vs stored, field by field, before any batch runs.
- Unknown stays unknown. The eligibility engine returns met / unmet / unknown — it never converts absence of evidence into a verdict.
- Nothing publishes on one measurement. Machine-detected signals are re-verified by an independent second method (different baseline statistic, country-normalization, duplicate-collapsed recounts). What fails is dropped and the drop is logged.
- We audit ourselves in public. Our own bugs become articles; the pipeline’s pulse is a public page.
3 · What the model can do today
- Find — every open tender in 27 countries, searchable in your language, including the 28K+ below-threshold contracts no aggregator lists.
- Qualify — buyer selection criteria structured for 87% of open above-threshold tenders; deterministic eligibility verdicts against a firm profile.
- Contextualize — for any tender: the buyer’s history, the incumbent field, what was paid before.
- Predict — 54–62% of buyer×category pairs repeat year over year (measured); the buyer-cycle layer turns that into a forward calendar.
- Detect — daily trending signals with historical analogues, independently fact-checked before publication.
- Investigate — threshold-gaming ratios, event correlations (floods, wars, budget cycles), concentration patterns — the Daily Tender class of findings.
4 · What we deliberately don’t claim
- No win-probability scores in the product. We built an experimental ML model; its honest benchmark is far below what we’d sell. Until a model beats deterministic context by a margin we can publish, it stays in the lab. (Vendors quoting impressive win-odds rarely publish their evaluation design. Ask them.)
- No causation from correlation. Our event pieces claim timing, not cause — the method box on each says exactly what was measured.
- No silent coverage claims. Sources publish differently: some retrospective-only, some missing categories. Where coverage is thin, our figures say so — they are floors, not ceilings.
- Translation is an index, not a legal text. Original titles are preserved everywhere; bids are won on source documents.
The asset behind this method: the data page. The method applied daily: The Daily Tender.