
For decades, Moore’s Law defined the pace of computing, setting the standard for how quickly performance improved while costs fell. That single metric reshaped the modern world, but the artificial intelligence sector now requires a different way of measuring progress. AI involves a complex mix of hardware, software, and data, meaning the industry needs a new lens to understand raw capability. A cheap token does not always mean cheap AI or a successful outcome. The true metric is the cost of finishing a useful job at an acceptable quality level.
Why Moore’s Law Is No Longer Enough
Moore’s Law has not ended, but it no longer tells the complete story. Leading-edge chips keep getting better, yet performance also comes from specialized architectures, memory, networking, and the software sitting above them. Silicon improves in steady generational steps. Model architecture improves in jumps when engineers change how the software works. A single blended rate averages across both, which describes neither accurately. The better lens is the price of a completed task that meets the required standard. Hardware affects how efficiently a model runs. Model design affects the computation required. Serving software, data, security, and human review determine if the result is usable or scalable.
Leading-edge chips continue to improve, while performance also comes from specialized architectures, memory, networking, and the software above them. It does not move as one curve. The cost of AI inference at a constant level of model performance fell roughly tenfold a year between 2021 and 2024, a factor of 1,000 in three years, according to Andreessen Horowitz. MIT researchers put the more recent frontier rate at five to 10 times annually. That rate is faster than transistor scaling ever delivered, and unlike Moore’s Law, it is a composite of several technological curves compounding rather than one.
From Token Prices to Qualified Outcomes
Cost per token alone does not tell an investor what useful work costs. A cheaper model may require more attempts or more human review. A more expensive model may complete the same task correctly in one pass. Some tasks are deterministic: The code passes, the numbers reconcile, or the required field is present. Others span chains of dependencies, several parties, and real-world interactions. Non-deterministic means the same process may not produce the same result each time, and reasonable people may prefer different outcomes. The modern economy handles that ambiguity through standards, review, and accountability.
Artificial Analysis offers one of the best current views of the trade-off. Its index covers agentic work, coding, scientific reasoning, and general knowledge, and reports average cost per benchmark task. Claude Opus 5 scores 61 at $2.03 per task. DeepSeek V4 Flash scores 44 at four cents. At the end of the day, a lower-scoring model may be the economic choice when it meets the standard requirements, and thus, we see deflationary impact on the cost of “intelligence.” The limitation is the word “task.” Benchmarks grade work inside a controlled environment. Valuable business processes often cross software systems, organizations, and the physical world, or depend on subjective judgment. Cost per task is a useful frontier, but not a universal price tag.
As the cost of a qualified outcome falls, more work becomes economical to automate or augment. Agents may consume more tokens as they plan, use tools, and recover from errors. This even expands into the physical area, where agents can summon or interact on behalf of individuals and organizations in the real world, or even act as embodied agents in real life. This amplification of scope, with access to systems and data, increases cybersecurity requirements. Agents running autonomously require watching observability, at a scale that dwarfs previous monitoring needs. The opportunity runs from chips and cloud infrastructure to networking, security, data platforms, business processes, and industry-specific applications.
Nebius Group (NBIS), an AI cloud provider, illustrates the enabling layer. Its Nvidia partnership spans AI-factory design, inference software, hardware deployment, and fleet management. It was also an early launch partner for Kimi K3, an example of the ecosystem benefiting when open-source models win. THNQ captures several layers pushing this frontier and commercializing the resulting capabilities. Moore’s Law taught investors to watch one curve. AI requires watching how chips, system architecture, models, data, security, and deployment improve together. Companies lowering the cost of useful work, or expanding what AI can do, are building the next phase of the market.
System Architecture and Software Innovation
Progress in this new metric extends beyond raw silicon to the detailed design of systems and software. Last week, at AMD’s Advancing AI 2026 event, AMD and Cerebras announced plans for a new inference system that pairs AMD’s Helios platform with the Cerebras Wafer-Scale Engine in a single workflow. AMD handles prompt processing and high-volume throughput, while Cerebras generates the answer. Company modeling projects up to five times more tokens per second per watt than a Cerebras-only configuration. The announcement shows how system architecture can add another efficiency curve.
Meanwhile, this week, China AI company Moonshot AI released Kimi K3, a new open-weight model that shows the software layer moving independently. It activates only 16 of 896 specialist components for each token it processes. Moonshot estimates that the design improves overall scaling efficiency by about 2.5 times over Kimi K2. Going deeper into the application and deployment layers, Nvidia also launched the Open Secure AI Alliance this week, bringing together companies across cloud computing, cybersecurity, enterprise software, and AI research. The alliance treats identity, permissions, guardrails, logs, and evaluation as part of agent security. It extends the coalition strategy Nvidia began at the model layer with the Nemotron Coalition, which pools research, data, evaluations, and computation across AI labs.
Leave a Reply