Public aliases. Each links to the full write-up.
Asked to build a per-request model router, I measured the fleet first and found model choice was about 1% of the bill.
A benchmarking platform where getting the number is the easy half. The hard half is deciding whether it is allowed to mean anything.
An analyst names a party; six sanctions regimes, adverse media, aliases and corporate ownership come back as a cited, defensible risk determination.
Planners ask supply-chain questions in plain English; 32 domain agents answer with SQL-grounded results, under role-scoped access the model cannot talk its way past.
Thousands of daily freight transactions scored for over- and undercharge by three deliberately asymmetric detectors, and an evaluation that turned out to be measuring the review interface rather than the model.
A hybrid surveillance system where a sequence model ranks alerts and a rules engine still decides what an alert is, because in anti-money-laundering, the part you cannot explain to an examiner is the part you cannot ship.
Models that predicted which accounts would roll to charge-off early enough to do something about it, built just as a payment-holiday programme made the word 'delinquent' mean something different.
Carrier invoices arrive as PDFs and the system of record holds what the shipment should have cost. Reconciling the two across millions of records turned an audit that had always been a sample into one that was complete.
Repossessed vehicles lose value every day they sit. A model that ranked recovery routes by expected net proceeds improved recovery profits 30%, and the reporting automation around it is where I learned what makes an analysis get used.
A weekly pipeline collapses a multi-million-row inventory history into one row per item and site, so an agent can answer any history question with a single small lookup.
A batch runtime across four model providers whose rate limits are discovered by a control loop rather than configured, and whose credentials never touch a developer's machine.
A weekly pipeline over planning, demand and lane-rate data that ranks mode-shift candidates by modelled freight savings and carbon impact.
Per-segment transit-time prediction by lane and shipping method, trained on eighteen months of closed orders and scored against the live backlog, with a validation gate that fails the pipeline if a join silently multiplied the rows.
Batch screening of the customer base against sanctions lists, where the whole design problem is deciding which side of the cross-join is a delta.
Validates carrier on-time-delivery root-cause codes against a declarative decision tree, with deterministic routing before the model and a human ground-truth workbook behind the evaluation.
A shipment-tracking assistant rebuilt from a single-file prototype into the same platform architecture as the enterprise agent estate, including a middleware guard that strips SQL the prompt already forbade.
Replaces a per-pallet rule of thumb with a physical packing model at part-number level and a learned correction from thirty-one months of actuals, cutting mean absolute percentage error from 31.9% to 6.0%.