Bank of America just launched an AI tracking tool. The code didn’t lie—this isn’t just another research report. It’s a systematic tracker covering model intelligence and costs, and for the crypto AI crowd, this is a seismic shift in how we evaluate on-chain AI tokens.
We didn’t see this coming from a traditional bank. But the timing is perfect. The crypto AI market is drowning in noise—every L1, L2, and rollup claims to be the next AGI. Yet there’s no standardized way to compare model performance vs. cost. BofA just filled that gap.
Hook On March 15, 2025, Bank of America’s global research team unveiled a new tool called the “AI Model Intelligence & Cost Tracker.” The announcement came via a cryptic tweet from their head of research: “We’re done with hype. Let’s track the real numbers.” No official press release, no fancy launch event. Just a login portal for institutional clients and a short document describing the methodology.
Context The tool aggregates two key metrics: model intelligence (benchmark scores from MMLU, HumanEval, MATH, etc.) and cost (API pricing per million tokens, inference compute, and estimated training costs). It covers 47 models so far, including GPT-4o, Claude 3.5, Gemini Ultra, Llama 3, and even select open-source Chinese models like DeepSeek-V3. The data is updated bi-weekly, with a feedback loop from BofA’s own trading desk analysts.
Why now? Because the AI arms race is a cost war. Every crypto project building an AI agent—from Render Network to Akash to Bittensor—needs to justify its token economics. If a model costs 10x more but delivers only 2% better accuracy, that’s a red flag. BofA’s tool quantifies that trade-off.

Core Based on my experience auditing DeFi protocols and tracking on-chain AI models, this tool is a methodological breakthrough. But let’s get into the technical details.
The tracker uses a weighted composite score: 60% model intelligence (benchmark performance), 30% cost efficiency (USD per million tokens), and 10% community adoption (Github stars, API usage growth). The intelligence score is not a simple average—it’s a rolling window that penalizes models that degrade over time (e.g., after fine-tuning with bad data).

I’ve seen similar attempts from places like LMArena and Artificial Analysis, but they’re all community-driven and lack financial rigor. BofA brings something new: a proprietary risk-adjusted model score. They apply a volatility discount to models that have shown inconsistent performance across different benchmarks. For example, if a model scores 95% on MMLU but only 60% on HumanEval, its composite gets penalized by 15%. This is exactly what crypto investors need when evaluating AI tokens—because a model that only works on curated benchmarks is a rug pull waiting to happen.
Here’s the kicker: the cost metric includes not just API calls but also the implied cost of data privacy. If a model requires sending data to a centralized server, BofA adds a “privacy premium” of 30% to the cost. This is a direct nod to the value of on-chain AI models that run locally or via decentralized inference networks. The data didn’t flinch, but the implications are huge: decentralized AI models like those on Bittensor or Akash could suddenly look more cost-effective on paper, even if their raw intelligence is lower.
But there’s a catch. The tool currently only covers models that are “commercially available via API or app.” That excludes many open-source models running on decentralized infrastructure. For example, the Llama 3 variants used by Render Network are not listed unless they have a public API endpoint. This creates a blind spot—the very models that crypto AI projects rely on are invisible to BofA’s tracker.
Contrarian View Everyone is celebrating this as a win for transparency. But I see a darker angle: BofA is positioning itself as the gatekeeper of AI evaluation. The tool’s methodology is proprietary, and the weights are not fully disclosed. This is exactly the same playbook as Moody’s or S&P—create a rating system, then sell access to it. In crypto, we’ve seen how centralized rating agencies can manipulate markets (remember the 2008 CDO ratings?).
Also, the tool ignores the most important factor for crypto AI: decentralization. A model’s score doesn’t account for whether it’s run on a permissionless network. A centralized model like GPT-4o might score high, but it can be shut down or censored at any time. BofA’s tool doesn’t penalize that risk. The market didn’t realize this yet, but it will—once a major AI token gets downgraded because its model is too centralized.

Furthermore, the tool’s reliance on public benchmarks is a known weakness. Models can “overfit” to these benchmarks, scoring high while failing in real-world tasks. For example, the recent drama with OpenAI’s GPT-4o scoring 98% on MMLU but failing basic math in a live demo. BofA’s tool doesn’t catch that—it only updates every two weeks. In crypto, where AI models can be fine-tuned and deployed in hours, that lag is deadly.
Takeaway Bank of America’s AI tracker is a double-edged sword. For crypto AI investors, it’s the first credible tool to compare models on a level playing field—but only if you trust the methodology. The real question is: will BofA open-source the tracker or keep it a black box? If they choose the latter, we’ll need a decentralized alternative built on-chain, where the weights are transparent and the data is immutable. Until then, use this tool as a starting point, not a final verdict. The code didn’t lie, but the code is hidden.