It feels like every three to five years, there is a new variant of the same pitch: you have everything under your eyes; you just need a new, fancy pair of glasses, and the truth will be revealed. We are only one stretch of additional compute away from the Truth, with a capital T.

The recent computing arms race across multiple industries is seductive. This is not without merit: increased compute has undeniably expanded what is technically feasible across modeling and execution. As every market participant ramps up their computing power, promises of spectacular bottom-line impact materialize. In finance, the number of GPUs is sometimes used as a proxy for team strength, as if more digital workers sieving market data round-the-clock guarantees more alpha discovered, faster, with minimal change to headcount or infrastructure. Compute, or more precisely, tokens, can be bought. It is as if LLMs risk becoming to quantitative research what dropshipping became to parts of online retail: highly accessible, but potentially difficult to differentiate.

But here is the inconvenient question: can insights truly be exported? And more crucially, what if there is simply no more alpha to be found in the data we already have?

These questions have often received less attention from vendors and many quants alike. Yet our current understanding of information theory tells a very different story. Information, by construction, is finite. There is no amount of preprocessing, post-processing, or transformation that will create genuine information from a dataset. What it creates, instead, is noise.

The objection is predictable: what if we have not fully maximized the information already in our datasets? What if there were a goldmine we had been sitting on all that time, unaware?

Theory and practice deliver a disappointing answer. The Pareto principle, an empirical rule well-verified in practice, holds that 80% of effects derive from 20% of inputs. In finance, this means that only a small fraction of our data contributes to the vast majority of achievable alpha. The critical 20% has likely been extensively explored, meaning most likely narrower pockets of opportunity.

Let's step back and look at the history of financial data, which tells a clear story. Computerization proved transformative; it dramatically improved data quality and reliability. By the 2010s, conventional wisdom already held that most alpha from purely financial datasets had been extracted.

Previous generations understood markets well. They did lack our modern tools, but the sequence of digitization of financial data proceeded from the most critical datasets to the least.

Alternative datasets gained adoption precisely because conventional data had been thoroughly mined. Those datasets have now been consolidated, categorized, and distributed to most market participants. A relatively small set of providers now supply increasingly standardized datasets.

Given the substantial research effort dedicated to alpha generation over decades, many of the most accessible insights have likely surfaced by now. If the low-hanging fruit were not harvested, then a faster machine will not find it now. More compute does not change this reality.

None of this diminishes the very real gains LLMs have delivered in productivity, prototyping, and research acceleration. But what GenAI will unfortunately also accelerate, is overfitting; and, more critically, alpha decay.

Understanding markets requires more than characterizing patterns. It requires grasping why relationships exist, their underlying conditions, and their chain of causality. Knowledge of patterns is not understanding. A model can pick up correlation; it cannot explain causation. Without the why, you are trading noise, not signal. This distinction matters more now than ever.

LLMs are trained on all publicly available data, so they can, in many settings, produce similar answers to the same questions, leading to partial convergence in reasoning pathways. Two funds relying heavily on LLM-driven models for alpha research, both utilizing similar datasets, may "discover" overlapping signals simultaneously, whether alpha or noise. We have automated the alpha convergence already observed in macro during the late 2010s, when portfolio managers and trading desk heads rotated across buy and sell sides, seeding identical ideas and reaping diminishing rewards from a saturated market.

In times of volatility bursts and steep regime shifts, the convergence of assumed uncorrelated strategies becomes a severe handicap. The catch is, if everyone is trading the same "discovered" signals, diversification fails. Speed may matter; understanding matters more. The past does not neatly explain the present, and the present does not reliably predict the future. Compute processes what has been; it cannot guarantee the right trades for what comes next.

This leaves three genuine sources of alpha. The first one is extending the underlying dataset; this can be done by creating new financial markets, or identifying ones where data continues to be sparse and inefficiencies persist. The second one is leveraging structural advantages, including privileged access to information, to structure trades. Lastly, the third one, which is becoming increasingly scarce in an age of computing power arms races, is genuine originality and independence of thought.

The fallacy of compute is not that processing power is useless. However, it lies in the narrative that more hardware can manufacture a signal, an alpha, that does not exist. Compute reliably optimizes execution; whether it leads to genuine discovery depends on how it is used.

Maybe, we should wonder, if in a world where everyone is racing to buy more chips, the real opportunity lies in stepping off the treadmill. And asking better questions.