Skip to content
Kiran

Résumé

Kiran

AI/ML engineering and technical leadership

I architect production AI systems end to end. Then I measure them honestly, even when that kills the brief.

Experience

Cisco

2023 – present

AI/ML engineering and technical leadership, global supply chain

Production ML → ML platforms → GenAI and RAG → agentic AI → enterprise AI architecture → measurement and cost.

  • Architected an enterprise natural-language-to-SQL agent platform end to end: 32 domain agents over SAP HANA, three surfaces, layered security, and a model-tier system operations can retune without a deploy. Other engineers built domains and surfaces against the contracts I set.
  • Built a multi-agent compliance investigation platform: an orchestrator and five specialists under one deadline, one source registry and four budgets, with fail-static authorization and durable checkpoints.
  • Rebuilt an inherited benchmarking tool into a measurement platform with validity gates, paired statistics, judge-reliability sampling and a written retraction policy.
  • Ran the fleet cost study that reframed a model-routing brief: measured the bill, found model choice was worth about one percent, and recommended budgets and cache hygiene instead.
  • Production ML across freight pricing anomaly detection, transit-time prediction, mode-shift optimisation, inventory statistics and warehouse capacity estimation. Each one was measured against the system it replaced.
  • Wrote and published an internal multi-provider LLM batch runtime that other teams adopted and vendored into their own projects.

Basis for this sectioncode

Deloitte

Aug 2021 – Dec 2023

Senior Data Scientist, financial crime and recovery analytics

Rules-based surveillance → hybrid rules-and-ML detection at transaction scale.

  • Strengthened anti-money-laundering surveillance over 10M+ transactions, improving detection accuracy 25%, which meant higher-risk alert yield with less investigator noise.
  • Deployed RNN transaction profiling at 74% suspicious-activity precision with a 30% false-positive reduction, lowering cost-to-investigate per alert on high-volume feeds.
  • Co-built an enterprise case-management system that cut alert processing time 40%, freeing investigator capacity for escalations rather than triage.
  • Automated insurance benefits recovery on AWS SageMaker, surfacing $2M in customer remediation and 40% faster claims review, with defective claims routed to manual review at 78% accuracy.
  • Reconciled transport-management against invoice data by parsing PDFs across millions of freight records: 100% discrepancy identification and $2.5M annual savings for a logistics client.

Basis for this sectionattested

BMW Financial Services

Jun 2020 – Dec 2020

Data Scientist, credit risk

Portfolio risk modelling under a regime shift nobody had training data for.

  • Built delinquency and charge-off models at 87% accuracy and 81% recall that reduced portfolio loss 20% through earlier collections targeting.
  • Delivered three production models above 85% accuracy for COVID-era 30/45-day delinquency segmentation, prioritising outreach before charge-off on heavily imbalanced portfolios.
  • Presented risk dashboards to the Executive Committee, including what the numbers did not yet support.

Basis for this sectionattested

Accenture (BMW Group)

Jul 2016 – Mar 2019

Senior Data Analyst, recovery analytics and enterprise reporting

Reporting automation → the first models I put in front of a business decision.

  • Designed a recovery model that improved vehicle recovery profits 30%.
  • Optimised Informatica ETL and SQL workflows, cutting batch runtime 25% for group reporting.
  • Automated reporting that saved 5+ analyst hours per week; recognised as Best Employee of the Team twice.

Basis for this sectionattested

Systems

Public aliases. Each links to the full write-up.

  • Asked to build a per-request model router, I measured the fleet first and found model choice was about 1% of the bill.

  • A benchmarking platform where getting the number is the easy half. The hard half is deciding whether it is allowed to mean anything.

  • An analyst names a party; six sanctions regimes, adverse media, aliases and corporate ownership come back as a cited, defensible risk determination.

  • Planners ask supply-chain questions in plain English; 32 domain agents answer with SQL-grounded results, under role-scoped access the model cannot talk its way past.

  • Thousands of daily freight transactions scored for over- and undercharge by three deliberately asymmetric detectors, and an evaluation that turned out to be measuring the review interface rather than the model.

  • A hybrid surveillance system where a sequence model ranks alerts and a rules engine still decides what an alert is, because in anti-money-laundering, the part you cannot explain to an examiner is the part you cannot ship.

  • Models that predicted which accounts would roll to charge-off early enough to do something about it, built just as a payment-holiday programme made the word 'delinquent' mean something different.

  • Carrier invoices arrive as PDFs and the system of record holds what the shipment should have cost. Reconciling the two across millions of records turned an audit that had always been a sample into one that was complete.

  • Repossessed vehicles lose value every day they sit. A model that ranked recovery routes by expected net proceeds improved recovery profits 30%, and the reporting automation around it is where I learned what makes an analysis get used.

  • A weekly pipeline collapses a multi-million-row inventory history into one row per item and site, so an agent can answer any history question with a single small lookup.

  • A batch runtime across four model providers whose rate limits are discovered by a control loop rather than configured, and whose credentials never touch a developer's machine.

  • A weekly pipeline over planning, demand and lane-rate data that ranks mode-shift candidates by modelled freight savings and carbon impact.

  • Per-segment transit-time prediction by lane and shipping method, trained on eighteen months of closed orders and scored against the live backlog, with a validation gate that fails the pipeline if a join silently multiplied the rows.

  • Batch screening of the customer base against sanctions lists, where the whole design problem is deciding which side of the cross-join is a delta.

  • Validates carrier on-time-delivery root-cause codes against a declarative decision tree, with deterministic routing before the model and a human ground-truth workbook behind the evaluation.

  • A shipment-tracking assistant rebuilt from a single-file prototype into the same platform architecture as the enterprise agent estate, including a middleware guard that strips SQL the prompt already forbade.

  • Replaces a per-pallet rule of thumb with a physical packing model at part-number level and a learned correction from thirty-one months of actuals, cutting mean absolute percentage error from 31.9% to 6.0%.

Capabilities

Agentic AI in production
Budgets at the tool-call boundary, wall-clock deadlines that survive a human-approval pause, durable checkpoints on evictable infrastructure, cross-instance run control, and authorization that fails static rather than open. Every one of those exists in my systems because something specific broke first.
Measurement and evaluation
Validity gates that derive what a report may and may not claim, paired statistics with an effect-size floor, judge-reliability sampling, a written retraction policy, and audits of the instruments themselves. Twice, auditing my own safety net found it was measuring nothing.
AI platforms and economics
Model routing as an operational lever rather than a code change, cache economics measured rather than assumed, blast-radius floors, and offline pre-computation that makes an expensive question cheap before an agent ever asks it.
Machine learning
Hybrid ensembles with deliberately asymmetric detectors, hierarchical fallback when the precise comparison group runs out of data, grain-and-fanout gates between extraction and modelling, and baselines implemented in code so the improvement is a comparison rather than an assertion.

A note on the numbers

Everything quantified on this site is labelled with how it is known: verified in code, observed in traces, calculated from published prices, modelled as a range, or attested by me or the business. Modelled figures are given as ranges and never as results. If a figure here matters to a decision you are making, ask me which of those it is and I will tell you the same thing the label says.

A printable PDF is not on this page yet; printing the page produces a readable copy.