PerfInsights: GenAI-powered detection of Go performance antipatterns at Uber
Optimizing Go services at Uber required deep expertise and significant manual effort, with profiling and analysis taking days to weeks; in March 2024, the top 10 Go services alone accounted for more than multi-million dollars in compute spend, making performance tuning prohibitively expensive and non-trivial for most teams.
Initial single-shot LLM-based antipattern detection produced inconsistent and unreliable results—responses varied between runs, included hallucinations, and often generated non-runnable code—with false positives exceeding 80%.
source quote
source quote
source quote
source quote
source quote
source quote
source quote
source quote
source quote
PerfInsights reduced performance analysis from days to hours, cut engineering time per issue from 14.5 hours to almost 1 hour (93.10% savings), reduced false positives from over 80% to the low teens, and produced hundreds of merged diffs driving compute cost reductions across Uber's Go services.
Show all 21 reported metrics
Compare to your context
Tell us your scale, team, and constraints. We'll show what changes at your size, what fails at your scale, and whether this case is a fit, needs adaptation, or won't scale to you. Free demo, no signup.
Frequently asked questions
What did this team achieve with this AI workflow?
PerfInsights reduced performance analysis from days to hours, cut engineering time per issue from 14.5 hours to almost 1 hour (93.10% savings), reduced false positives from over 80% to the low teens, and produced hund…
What tools did this team use?
PerfInsights, LLM, LLMCheck, Optix.
What results were reported?
compute spend (top 10 Go services, March 2024): more than multi-million dollars; Analysis time reduction: tasks that once required days now take hours; False positive rate (after validation pipeline): over 80% to the low teens; Diffs generated and merged: hundreds (source-reported, not independently verified).
What failed first in this deployment?
Initial single-shot LLM-based antipattern detection produced inconsistent and unreliable results—responses varied between runs, included hallucinations, and often generated non-runnable code—with false positives excee…
How is this quality assurance AI workflow structured?
Fleet-wide production profiling → Hotpath function filtering → Static noise exclusion → LLM antipattern detection → LLM jury validation → LLMCheck rule-based validation → Confidence-scored output → Integration with Optix → Accuracy tracking feedback loop.
Related quality assurance cases
Now compare it to your context
This case is one data point. Whether its pattern fits you depends on your volumes, your stack, and your exception load — that comparison is the step no case study can do for you.