← Back to portfolioArchitecture decisions log

The load-bearing decisions, and the alternatives we rejected.

This log captures the architecture choices that shape how every module behaves. Each entry follows the same shape: the context, the decision, its consequences, and what we considered and rejected. Entries marked revisiting are ones we expect to reopen.

DEC-0012025-10Accepted

Deterministic extraction over LLM extraction for financial data

Context
Liasses fiscales, grand livre, and FEC files carry the primary financial signal for every deal. A single wrong EBITDA propagates into the valuation, the IM, and every buyer conversation. Cost of an error is high and hard to detect after the fact.
Decision
Never let an LLM generate financial figures. PDF parsing runs through Docling (IBM, MIT-licensed) for deterministic table extraction. Excel parsing runs through openpyxl. The LLM is only permitted to classify documents and flag anomalies.
Consequences
We pay the cost of building a richer extraction pipeline (form recognition, table normalisation, FEC schema handling) instead of outsourcing the work to a prompt. Extraction is reproducible and auditable. Edge cases require explicit rules rather than hoping the model notices.
Rejected alternative
A single prompt that ingests the whole PDF and returns structured JSON. Rejected because hallucination rates on numeric extraction at this document length are incompatible with the use case.
DEC-0022025-11Accepted

A single LLM gateway — no direct SDK imports in business modules

Context
Frontier model leaderboards change monthly. Being locked to one SDK means every model migration becomes a codebase refactor touching dozens of call sites. Price/quality tradeoffs also differ per task type (reasoning, structured output, web extraction).
Decision
All frontier calls go through a single client at infra/llm_client.py. Business modules import that client only. Model routing is config-driven: task type → model name. Switching from Claude to GPT to Gemini is a YAML change.
Consequences
Adds one layer of indirection and a small amount of per-request overhead. Forces us to standardise on a minimal common-denominator API surface. Vendor-specific features (Anthropic caching, OpenAI structured outputs) need explicit gateway support before any module can use them.
Rejected alternative
Per-module SDK imports for maximum flexibility. Rejected because the cost of one migration in the first three months outweighed the cumulative cost of the abstraction.
DEC-0032025-11Accepted

HTMX + Jinja over React for the cockpit

Context
The cockpit is an internal control surface for a two-person team. Five tabs (Sourcing, Deals, Buyers, Analyst HITL, VP view), all read-heavy, all server-rendered. Peak concurrent users: two.
Decision
FastAPI + HTMX + Jinja2 templates. Server-rendered HTML with partial updates. No SPA framework, no build step for the UI, no state management library.
Consequences
Shipped the first working cockpit in three weeks instead of eight. Deployment is a single container. Debugging is HTML in the DevTools. The cost is that if the product ever grows to an external user base with real interaction patterns, we'll need to port to React.
Rejected alternative
Next.js + React Query. Rejected for this cockpit; retained for the public portfolio site where we actually need SPA capabilities.
DEC-0042025-12Accepted

LangGraph only for the Email VP agent, not as a general orchestrator

Context
LangGraph is powerful but introduces real complexity: stateful graphs, checkpointing, debuggability tradeoffs. Most of our modules are linear pipelines (sourcing → enrichment → scoring) that do not need stateful orchestration.
Decision
Use LangGraph only where stateful memory across weeks matters — specifically, the Email VP agent that needs to remember that Argos asked for exclusivity on March 3rd when they reply on March 20th. Everywhere else, plain sequential Python with Pydantic contracts at the boundaries.
Consequences
Most of the codebase stays readable by anyone who knows Python. Onboarding a contractor on a business module does not require learning a framework. The VP agent's persistence is a solved problem via PostgreSQL checkpointing.
Rejected alternative
Wrapping every module in a LangGraph state machine. Rejected because CrewAI / AutoGen / full-LangGraph deployments add a framework tax without a matching benefit for linear pipelines.
DEC-0052026-01Accepted

Build the 7M-record buyer database without Pappers

Context
Pappers provides a clean API for French company data at a real cost per call. The raw signal it provides (INSEE identity, INPI filings, BODACC publications) is publicly available from the original administrations.
Decision
Ingest directly from INSEE (SIRENE stock files), INPI RNE (beneficial owners, directors), BODACC (acquisition history, legal notices), and open procurement data. Run our own entity resolution on top (see case study). Pay with engineering time, not per-call fees.
Consequences
The database becomes a defensible asset: we own the schema, the refresh cadence, and the enrichment layers. Every enrichment we add compounds on the next deal. Downside: we carry the maintenance burden when upstream schemas change, which they do.
Rejected alternative
Pappers-as-a-service. Rejected because at our expected query volume, the cumulative API cost would exceed the engineering cost of the in-house build within the first year, and we would hold no IP at the end.
DEC-0062026-02Accepted

No Send Ever — not now, not in V2, not ever

Context
Every generative system faces the question of how aggressive to be with autonomous external action. In M&A, an agent sending an email to a buyer without banker review is a business-ending risk: confidentiality breach, mistimed outreach, tone that misrepresents the mandate.
Decision
No agent may send any external communication. Ever. The HITL queue in the cockpit is the terminal stop for every agent output. The banker reviews, edits, or rejects, and the banker presses send. This is a product invariant, not a configuration.
Consequences
Throughput is capped by banker review time. We accept this because the reputation cost of one misfired email exceeds the throughput gain of any automation. We have received requests to relax this rule and have refused every one.
Rejected alternative
'Auto-send with human override within 15 minutes' policies. Rejected because the failure mode is adversarial (a bad email reaches a buyer before the override fires) and cannot be undone.
DEC-0072026-03Accepted

Hybrid RAG (dense + BM25 + RRF + cross-encoder) over embed-only

Context
Pure embedding-based retrieval under-recalls exact numerics and named entities. In M&A retrieval, the query is often 'what was the 2023 EBITDA for this specific company' — a question where approximate semantic match is worse than useless.
Decision
All retrieval uses Qdrant with hybrid search: bge-m3 dense embeddings + BM25 sparse, fused via Reciprocal Rank Fusion, then reranked with bge-reranker-v2-m3 cross-encoder. Namespaces are strict: deal_{id}_{type}.
Consequences
Retrieval quality on numeric and named-entity queries is materially better than embed-only baselines. Latency budget per retrieval increases to ~200–400ms. Qdrant hosting and cross-encoder inference add infrastructure cost. Worth it for the correctness gain.
Rejected alternative
Embed-only with a single dense model. Rejected after measured recall on a golden set of M&A queries was unacceptable for financial figures.