Friedrich & ClementBETA
How we know what we know

The model, the method, and its limits.

Everything we publish — every tender page, every signal, every Daily Tender figure — comes out of one data model and one verification culture. This page is the whole of it.

1 · The model: a graph of who buys what from whom

The corpus is not a pile of notices; it is a connected graph built from them:

  • 18.6M notices — the events: who asked for what, when, under which procedure, above or below the EU threshold.
  • 13.4M awards — the outcomes: winner, price, date, linked back to their notices.
  • 1.6M companies — the winners as entities, 1.2M anchored to official registry identifiers (NIF, SIREN, CUI, EIK, Y-tunnus…) so namesakes never merge and subsidiaries never blur.
  • Normalized buyers × CPV categories × years — every buyer’s purchasing rhythm, machine-readable.

Most procurement data stops at the first bullet. The intelligence lives in the joins: the same records, connected, answer questions the records alone cannot.

2 · The method: gates, not vibes

The pipeline was rebuilt around hard gates after we caught it silently losing fields at scale in its early months. The rules since:

  • Never guess a field. Every source integration starts from the portal’s real responses; every field is mapped or explicitly skipped with a reason.
  • One record end-to-end before a million. Dry-runs compare source vs stored, field by field, before any batch runs.
  • Unknown stays unknown. The eligibility engine returns met / unmet / unknown — it never converts absence of evidence into a verdict.
  • Nothing publishes on one measurement. Machine-detected signals are re-verified by an independent second method (different baseline statistic, country-normalization, duplicate-collapsed recounts). What fails is dropped and the drop is logged.
  • We audit ourselves in public. Our own bugs become articles; the pipeline’s pulse is a public page.

3 · What the model can do today

  • Find — every open tender in 27 countries, searchable in your language, including the 28K+ below-threshold contracts no aggregator lists.
  • Qualify — buyer selection criteria structured for 87% of open above-threshold tenders; deterministic eligibility verdicts against a firm profile.
  • Contextualize — for any tender: the buyer’s history, the incumbent field, what was paid before.
  • Predict — 54–62% of buyer×category pairs repeat year over year (measured); the buyer-cycle layer turns that into a forward calendar.
  • Detect — daily trending signals with historical analogues, independently fact-checked before publication.
  • Investigate — threshold-gaming ratios, event correlations (floods, wars, budget cycles), concentration patterns — the Daily Tender class of findings.

4 · What we deliberately don’t claim

  • No win-probability scores in the product. We built an experimental ML model; its honest benchmark is far below what we’d sell. Until a model beats deterministic context by a margin we can publish, it stays in the lab. (Vendors quoting impressive win-odds rarely publish their evaluation design. Ask them.)
  • No causation from correlation. Our event pieces claim timing, not cause — the method box on each says exactly what was measured.
  • No silent coverage claims. Sources publish differently: some retrospective-only, some missing categories. Where coverage is thin, our figures say so — they are floors, not ceilings.
  • Translation is an index, not a legal text. Original titles are preserved everywhere; bids are won on source documents.

The asset behind this method: the data page. The method applied daily: The Daily Tender.