Some time ago, a critical vulnerability went unnoticed in LiteLLM, one of the most widely used libraries for connecting applications to language models like GPT or Claude. It wasn't a minor bug. It was a backdoor that allowed arbitrary code execution on servers, with potential access to model credentials, data in transit, and more. What was revealing wasn't the technical flaw itself. It was how long it took to be detected, the massive dependency on that package, and the fact that many organizations—some making decisions with real impact on people's lives—were completely unaware of its existence.

It's worth examining not the bug itself, but the environment that allowed it to thrive.

For those who don't work in tech: LiteLLM acts as a universal adapter. It connects applications to AI providers like OpenAI, Anthropic, or Google through a common language. It's essential in healthcare systems, educational platforms, and financial analysis tools. That invisible adapter is rarely audited, even though everything depends on it.

The problem is structural opacity. Bugs arise in any software; they're an inevitable part of complex structures. What's critical is that LiteLLM, being open source and theoretically auditable, is rarely audited in practice. Modern applications depend on hundreds of packages, each with its own risks and maintainers. Few organizations review them rigorously. The promised transparency remains theoretical; most rely on reviews by others that often simply don't happen.

This dynamic repeats itself across different contexts. Critical infrastructure gets built quickly and adopted even faster, while auditing arrives late, if it arrives at all. Technological history offers concrete examples: the BGP protocol, the backbone of internet routing, was designed in the eighties assuming mutual trust between operators that no longer exists. The SWIFT interbank payment system failed in 2016 because banks lacked controls that had been taken for granted. These opaque structures concentrate power in the hands of those who control or understand them best, not always in those who should hold it.

What's unsettling is how this contradicts the discourse around AI governance. We hear talk of "auditable AI," explainable systems, and algorithmic accountability. These are valid goals. But if we don't actually audit, in practice, the packages that link applications to models, what's the point of auditing the AI that approves loans, recommends parole, or prioritizes medical diagnoses? Trust breaks down in the middleware, in that invisible dependency nobody bothered to inspect.

When analyzing AI in justice systems, bias in models turns out to be only part of the problem; the surrounding architecture lacks real verification. LiteLLM illustrates the same point. A transparent model doesn't save an opaque system if the connection to the real world isn't continuously verified.

This is where the historical perspective gains weight. In the nineteenth century, as railroads expanded as key infrastructure, power resided with those who controlled nodes, schedules, and fares—not with governments or users. The same happened with telecommunications networks, financial systems, and content algorithms. In emerging economies, that concentration of power exacerbated local inequalities; today the same process is being replicated in AI adoption without safeguards. Opacity isn't an accident: it emerges from rushed designs that never built in transparency from the start.

This is becoming urgent because AI adoption in governance—public resource management, risk assessment, service allocation—is moving faster than auditing mechanisms can keep up. This isn't an argument against AI; that would be illogical. It's a critique of institutional haste and the mistaken notion that open source automatically implies auditing. Importing complex tools into critical systems requires first building adequate verification.

There's no complete solution in sight. The problem is more intricate than it appears. But there are necessary steps: continuous, automated audits of dependencies as a requirement in critical sectors; technical teams within regulatory bodies with direct inspection authority; frameworks that cover the entire chain—model, middleware, data, infrastructure—not just the visible algorithm.

Building a society where AI orchestrates fair and verifiable collective decisions requires innovations that are still missing today. The LiteLLM hack goes beyond the technical: it reveals where we actually stand on this path, beyond the optimistic announcements.

Stones don't lie, but historians sometimes do.


Sources:

1. Trail of Bits / Socket Security — Analysis of vulnerabilities in software dependency chains (2023-2024)

2. SWIFT Banking System Breach — Bangladesh Bank heist, Federal Reserve Bank of New York reports (2016)

3. BGP Security History — NIST Special Publication 800-54, Border Gateway Protocol Security

4. Wired / Ars Technica — Coverage of the LiteLLM ecosystem and vulnerabilities in AI middleware (2024)

5. AI Now Institute — "Algorithmic Accountability Policy Toolkit" (2023)