← Back to portfolioCase study · Engineering

Resolving 500,000 French company entities per day at near-zero marginal API cost.

This is the unsexy work that makes the rest of the system possible. The architecture is a 3-tier cascade: deterministic matching first, local inference second, frontier model only as a last resort. What follows is how we arrived at those numbers, what each tier actually does, and the tradeoffs we consciously accepted.

Scale
~500K candidate entities / day
Cost (naive)
~$250K / day via frontier API
Cost (our cascade)
~$1–3K / day, dominated by fixed GPU depreciation
Failure rate @ tier 1
~30%
Escalation to tier 3
<2%
P99 latency
~4.2s (when tier 3 fires)

Same company, fifteen different spellings, six different portals.

Small-cap M&A sourcing means ingesting 10+ deal portals (BPI, Fusacq, CessionPME, Michel Simond, and seven others), each with its own freeform description of the target. The same company might appear as “Plomberie Dupont”, “Plomberie Dupont SAS”, “Ets Dupont Père & Fils”, “Dupont - Plombier à Lyon 69”, and “DUPONT (SIREN 123456789)”. Left unresolved, we’d pay multiple times for the same lead and pollute downstream signal with duplicates.

Resolution means mapping every candidate string to a canonical SIRENE record (the French statistical-institute national register). When SIRENE doesn’t have it, we escalate: INPI RNE, BODACC, open company-data aggregators, and eventually manual review.

LLM-only resolution: $250K/day, and wrong.

The obvious first instinct is to prompt a frontier model with a candidate name and ask it to return a canonical SIRENE ID. It works on a toy dataset. It fails at scale for three compounding reasons:

  • Cost. At roughly $0.50 per resolution with Claude or GPT-4 and the volume we process, the bill would reach six figures per day before the end of week one.
  • Latency. Each call is a 1–3 second round-trip. For 500K entities, the serial version takes ~10 days and the parallelised version hits provider rate limits within minutes.
  • Accuracy. Frontier models hallucinate SIREN numbers that don’t exist. In a domain where a wrong identifier cascades into the wrong filing, the wrong financials, and the wrong valuation, this is not acceptable — even at 99% accuracy, 1% hallucination on 500K entities is 5,000 poisoned records per day.

The constraints force a different architecture: do as little LLM work as possible, and verify everything deterministically.

Three tiers. Each one fires only if the previous fails.

Tier 1

Deterministic fuzzy match

~70% resolved · $0 · <50ms

  • Normalise strings (remove legal forms, accents, extra whitespace, trailing city/postal hints).
  • Probe against an indexed Postgres table of ~4M resolved SIRENE records with pg_trgm similarity.
  • If top match ≥ 0.92 similarity AND legal-form or postal-code match on the tie-breaker, accept.
  • Cost: disk seek plus index lookup. No GPU, no API. This tier dominates throughput.
Tier 2

Local Qwen 2.5 32B-AWQ on vLLM

~28% resolved · ~$0 marginal · ~600ms

  • Candidate retrieval: top-10 similar SIRENE records by pg_trgm + sector embedding similarity.
  • Qwen 32B (AWQ 4-bit, ~18GB VRAM on an RTX 4090) is prompted with the candidate string plus the 10 candidates and asked to pick one or return none.
  • Output is constrained JSON via vLLM guided decoding; the model cannot return a SIREN that isn’t in the shortlist.
  • Zero per-call API cost. Amortised fixed cost from the GPU and power draw.
Tier 3

Frontier model via gateway

<2% resolved · $0.50 / call · ~2.5s

  • Reserved for ambiguous cases where Qwen returned none or its confidence was below a calibrated threshold.
  • Expanded context: additional LinkedIn signals, BODACC filings, web snippets fetched via Kimi 2.5.
  • If this tier also fails, the entity goes to the HITL queue for human adjudication. This is the only path through which a resolution can be accepted without a match in an authoritative register.

What the cascade actually buys us.

DimensionLLM-onlyCascade
Cost per resolution~$0.50~$0.002 (blended)
Daily cost @ 500K~$250,000~$1,000
P50 latency~1.8s~45ms
P99 latency~4.0s~4.2s (when tier 3 fires)
Hallucinated IDs~1% of outputs0% (constrained decoding + register check)
Throughput ceilingAPI rate limitPostgres index + GPU batch

The headline number isn’t the cost reduction — it’s the hallucinated ID count going to zero. Tier 1 and Tier 2 cannot, structurally, return a SIREN that doesn’t exist: Tier 1 looks up against an index of real records, Tier 2 uses guided decoding against a shortlist drawn from the same index. Only Tier 3 has degrees of freedom, and its output is itself validated against SIRENE before acceptance.

What we gave up, and why we’re fine with it.

  • We cannot discover genuinely new entities. If a candidate company isn’t in SIRENE (e.g. it was just registered yesterday and our mirror is 48 hours old), Tier 1 misses, Tier 2 has no shortlist to pick from, and Tier 3 goes to HITL. That’s the right behaviour for our use case: M&A targets are, by definition, companies old enough to be acquirable.
  • Tier 2 depends on the quality of the shortlist.If our similarity retrieval misses the right candidate in its top-10, Qwen cannot rescue it. We tuned the similarity threshold to prefer recall over precision at the retrieval stage, at the cost of Qwen doing more work.
  • A single RTX 4090 is a throughput cap. Qwen 32B at AWQ 4-bit yields ~40–60 tokens/sec per request on a 4090; at our current volume this is not a bottleneck, but scaling past ~2M entities/day would require either batching improvements or a second GPU.
  • The cascade assumes hallucination is always bad.There are domains where a plausible-but-wrong guess is better than nothing. M&A is not one of them. The cascade enforces a hard “no invented identifiers” property at the architectural level, not at the prompt level.

The cascade is not specific to entity resolution.

The same shape shows up everywhere we touch batch LLM work: deterministic first, local second, frontier only if the reasoning justifies it. It surfaces in document classification (regex → Qwen → Claude), in financial restatement proposals (rules → Qwen → HITL), in buyer-scoring (SQL filters → Qwen ranking → frontier reranker).

The lesson we keep relearning is that the interesting question is never “which model is best at task X?” — it is what does the cheapest correct path look like? For most high-volume tasks, the answer has the first two tiers carrying the load.