(Almost) Everyone's an Empiricist Now
Mapping frontier AI neolabs by their theory of knowledge

A quick note before I begin — the neolabs I discuss here have acquired, in about two years, both that nickname and a small literature. I assume a passing acquaintance with both. For anyone who wants orientation: the WSJ's These Billion-Dollar AI Startups Have No Products, No Revenue and Eager Investors is a good introduction to the phenomenon, and NEA's The Neolab Wild West is the most complete map I know of, though keep in mind that since it is a venture firm's thesis, it sorts these labs by where financial value is likely to accrue — corpus, training loop, verifier; frontier generalist, sovereign, science, robotics.
My work here is to sort the many of those labs by the argument that each one posits about what 1) knowledge is, and 2) where it comes from. If the empiricism-rationalism dispute is a distant memory from an introductory course, the Stanford Encyclopedia's entry on the innateness of language will do most of the remedial work. Thanks as always.
Introduction
Frontier tech has no shortage of drama. Most people's exposure to it came, I think, via Sam Altman's ouster from OpenAI in November 2023, his reinstatement days later, and the texts from that weekend that surfaced years afterward.

You might also be familiar with the technical disagreements within these frontier labs/neolabs and the arguments they put forth concerning the composition of training data or the allocation of compute. But, underlying these technical disputes is, I will argue, a prior question about 1) what knowledge is, and 2) how a constructed system (a model) might come to possess it. On that question the neolabs divide along lines that epistemology has occupied since the seventeenth century. The contest that appears to be about engineering is, at a more fundamental level, a re-staging of the dispute between empiricism and rationalism over the origins of knowledge.
My goal with this essay, which is both longer and topically heavier than my usual writing, is to map the frontier neolabs on to which camps/fields I think that they occupy, and then to draw two consequences from this mapping. I will contend, first, that the most visible and most heavily capitalized dispute in the field is not, as one might assume, a confrontation between opposed philosophies of mind (empiricism vs. rationalism), but a disagreement conducted almost entirely within the empiricist camp. By contrast, the genuine rationalists, those who hold that no quantity of data can substitute for innate structure, are a marginal presence, both intellectually and by the amount of capital they have received in funding. Second, I contend that the neolabs share a more consequential omission: each advances a detailed account of how intelligence is to be produced while presupposing, and nowhere defending, an account of what understanding is. Where a theory of understanding is required, a performance metric is installed.
I should preface that I am not trying to argue that philosophical analysis settles the question of which neolab will succeed. History counsels against such confidence, and nothing within epistemology determines what will emerge from a training run. The conclusion that I arrive at is a more limited one – that the disagreements between the neolabs themselves are more coherent, and the alliances between rivals stranger. Overall, my hope is that by seeing these positions clearly, you and I and people far more capable than us can identify which bets look brave and which are confused.
The two axes
The intuitive way to organize this would be to distribute the labs between two classic positions: empiricism, on which knowledge is built from experience, and rationalism, on which reason is the primary source and test of knowledge, and at least some of the mind's principles are knowable a priori.
This partition will not be my primary instrument for two reasons. The first is that almost every neolab of consequence occupies a hybrid position, and as such, the binary dissolves at the points where the most interesting disagreements lie. The second is that the partition, in its general form, has already been made – the idea that large language models revive the seventeenth-century quarrel is by now a settled observation, with its own entry in the standard reference works. I am not in the business of diluting the work of those far more capable and intelligent than me.
What has not been done, and what I am attempting to do here, is the assignment of the specific neolabs to positions within that terrain.
So for all those reasons, I sort along two axes that the neolabs themselves contest, rather than along a distinction artificially imposed upon them. The first axis concerns the evidence to which a system or model is bound – whether it learns from text, from sensory and spatial data, from embodied interaction, from experience it generates for itself, or from a formal principle specified in advance.
The second concerns the quantity of structure that must be present before any such evidence can instruct the model at all, ranging from the view that structure should be held to a minimum and permitted to emerge, to the view that without the correct architecture no understanding is possible.
The first position: Scalers

Language as sufficient
The first and presently dominant position holds that the knowledge worth having and building upon is already contained in the recorded corpus of human language, and that a system trained to predict with sufficient accuracy and at sufficient scale will acquire understanding as a byproduct that requires no separate provision.
Its adherents are the established neolabs (OpenAI, Anthropic, xAI, Mistral, DeepSeek, most of the Chinese frontier) and its philosophical ancestor, whether or not those adherents would accept the attribution, is Hume. Hume's account of the mind reduces its operations to impressions and the associations formed among them, and causation, in his analysis, is inferred from the constant conjunction of one event with another. A system that is trained to predict the next element of a sequence from those preceding it performs a version of exactly Hume's inference, extracting from an immense corpus the regularities of what tends to follow what. The corpus stands in for a lifetime of impressions, and association is executed at a scale that Hume certainly could not have imagined.
This position's clearest statement is this essay by Richard Sutton, whose thesis is that "general methods that leverage computation are ultimately the most effective." The lesson Sutton draws is, by his own description, a bitter one, because it is anti-humanist in that it instructs the researcher to withhold the very intuitions about the structure of cognition that a career of study has made compelling on the grounds that systems which learn their structure from data have repeatedly overtaken those into which structure was deliberately built. This is fundamentally a statement of the empiricist creed. It counsels distrust of one's priors about what a mind requires and recommends, in their place, data and computation.
This position merits a distinction between the epistemology that its proponents profess and the epistemology that their practice reveals. Their operative standard is not, in fact, empiricist, but pragmatist in a strict sense, because to them, a model is judged superior when it attains a higher score on an established benchmark. As such, that truth is discharged as measured performance and understanding is deferred as a promissory note. This operative standard is perhaps nearer to William James than to Hume, but it is seldom stated.
The practice has diverged from the profession in a second way. By now the scaling labs are no longer text purists in method, whatever their founding rhetoric. Instead, their frontier models are shaped as much by reinforcement learning against verifiers, by agentic training in environments, and by synthetic data as by next-token prediction over a human corpus. The first position has, in other words, been absorbing the second position's methodology (which you will read about next) without revising its stated epistemology.
A critic might say this dates any map that files these labs under "language as sufficient." I suppose my response to that would be that a camp that adopts its rivals' methods whenever they raise benchmark scores, while leaving its story about knowledge untouched, is supplying further evidence that its real commitment is the pragmatist one. The professed epistemology is a founding myth; the revealed epistemology is whatever wins.
Lastly, It would be a mistake for me to treat this camp as homogeneous simply because its members share a premise. The distances between OpenAI, xAI, Mistral, and DeepSeek are present, but lie off the two axes I am using. They more so concern efficiency, openness, and distribution far more than any disagreement about where knowledge comes from. Anthropic is the more interesting exception, however, because while it scales as the others do, it pairs the scaling bet with a rather heavy investment in interpretability — the attempt to reverse-engineer what a trained model has learned. This could be interpreted potentially as a hedge against the position's own creed, because a lab that was fully confident that understanding falls outside the realm of prediction would have little reason to spend so much effort trying to recover it after the fact. So while Anthropic scales like a Scaler, it seems to be betting that prediction alone may not tell it what its models understand.
The second position: Experientialists

Experience in place of inheritance
The second position accepts the empiricist premise that intelligence is learned rather than built in, but rejects the recorded corpus as the data to learn from. Its charge against the first position is that a record of human language is a record of human achievement, and that a system trained to imitate it can, at the limit, only reconstitute what its makers already knew.
The remedy is to sever the system from human data entirely and to let it learn as an organism does, from the consequences of its own actions in an environment. The founding document for this position is a manifesto by David Silver and Richard Sutton, announcing what they call an "Era of Experience," whose central claim is that training on human data imposes a ceiling, since "agents cannot go beyond existing human knowledge," and that the route past that ceiling runs through rewards grounded in an environment's actual consequences rather than in a human evaluator's prior judgment. David Silver's new lab, Ineffable Intelligence, carries the thesis to its terminus: a system meant to acquire everything from its own experience, beginning from no human data at all, whose very name declares that the knowledge it reaches will lie beyond the reach of human words.
You may have noticed that Sutton anchors this both position and the previous one. This is because his "bitter lesson" has two halves. The first position takes the half that says learning at scale defeats hand-crafted knowledge. This second position takes the sharper half, which concerns the source of the learning signal itself – not a corpus of human artifacts, but the system's own interaction with a world that answers back. The philosophical lineage present here is something closer to the constructivism of Jean Piaget, on which knowledge is assembled by acting on an environment and registering what results (a radical, non-linguistic empiricism in which even the concepts are to be discovered rather than supplied – hence why he considered children "little-scientists"). The example that this position invokes is a game-playing system that reached superhuman strength from self-play alone, having been shown no human games, and the wager is that this template generalizes beyond the closed worlds of board games to the open one.
Within this camp (as with all) there is a disagreement about how much may be given in advance. To learn from experience rather than from human data is one commitment, but to insist, further, that even the structure doing the learning must itself be discovered is a second and more radical one. Not everyone who accepts the first accepts the second. A related bet, though it belongs at a slightly different point on the second axis, is being made by neolabs pursuing recursive self-improvement (such as the aptly named Recursive Superintelligence), so that the improvement loop closes without continuous human direction. The epistemology here is thinner than a theory of knowledge (more so a wager on a mechanism). Still, it shares this position's conviction that human knowledge is a floor to be departed from rather than a summit to be approached.
The third position: World-Modelers

The world as the condition of content
The third position also rejects the sufficiency of language, but for a different and more philosophical reason. Where the second position holds that text is a ceiling on capability, the third holds that text is inherently a defective source of meaning because it is a thin, downstream projection of a physical reality the model has never encountered, from which no amount of prediction can recover the thing projected. Its founding text predates the current expansion – a 2022 essay by Yann LeCun and Jacob Browning whose central claim is that a system "trained on language alone will never approximate human intelligence," because the greater part of human and all of animal knowledge is sensorimotor and tacit, and language encodes only a fraction of it. Fei-Fei Li advances the same argument in a spatial register, describing today's models as "eloquent but inexperienced, knowledgeable but ungrounded," and treating spatial structure as the scaffold on which cognition is built.
The philosophical lineage here is neither Hume nor Piaget but James Gibson's ecological psychology and the tradition of embodied cognition it informs, on which perception is organized for action and meaning is anchored in a body's traffic with a structured environment (a realist stance). When LeCun designs architecture that predicts within an abstract representation rather than reconstructing raw perceptual detail, he is implementing a philosophical thesis where understanding resides at the level of structure. Or when Li insists on world models constrained by physics and geometry, she is wagering that the format of real cognition is spatial far before it is verbal (very Gibson!!!). The neolabs in this camp differ over which slice of the non-linguistic world to privilege (three-dimensional environments, driving, gameplay, robotic interaction, or raw recorded action) but they are unified by a single term ("grounding") and by the conviction that this is precisely what a text-trained system lacks.
The unity of those in this position is at the level of principle. Having agreed that a system must be bound to the world, these labs then divide immediately over which world: the three-dimensional space of a rendered environment (World Labs), the raw video of a driving scene (Wayve), the recorded motion of a body (Physical Intelligence), the moment-to-moment action of a game or a screen (General Intuition and Standard Intelligence). Each of these is a bet about which slice of non-linguistic reality carries the structure that text lacks/discards, but they disagree downstream of a shared conviction.
The fourth position: Structuralists

Structure as a precondition of knowledge
The fourth position is the only one of the five that is truly, truly rationalist, and it is by some distance the smallest, whether you measured that adherents or by capital deployed. Its claim is that no arrangement of data, however vast or however richly sensory, can by itself yield understanding, because understanding requires structure that must be present in advance of learning.
The position has three notable variants. The first, associated with Gary Marcus and descending from Fodor and Pylyshyn's argument that thought is systematic and productive in ways statistics cannot explain, holds that intelligence requires compositional, symbol-manipulating architecture, and that systems lacking it will generalize unreliably however much they are scaled.
The second, developed by Karl Friston, derives the conditions of agency from a single formal principle: any system persisting in an uncertain environment must act to minimize its long-run surprise. This is openly Kantian in structure, treating perception as top-down inference from a prior model rather than passive registration of stimuli.
The third, Jeff Hawkins's theory of cortical reference frames, holds that the brain supplies coordinate systems, inherited from the machinery of spatial navigation, as the innate format for all knowledge (abstract knowledge included).
What unites these variants is a commitment that the other four positions reject in one form or another — that the mind (natural or artificially engineered) brings structure to experience instead of deriving structure from it. This is, in the plainest sense, Kant's correction of the empiricists, his insistence that the forms through which we apprehend a world are conditions of experience rather than deposits left by it, and his formula that "intuitions without concepts are blind" serves as an epigraph for the entire camp.
It is also worth noting that despite their shared commitment, these three variants would not recognize one another as allies. A symbolic architecture, a variational principle, and a cortical reference frame are all very different proposals about what the missing structure is, and their proponents disagree about the details. What places them in a single camp is more so a share conviction that the structure cannot simply be read off the data, however much of it there is.
That the position is so sparsely funded relative to the others is itself a fact in need of explanation. I return to it below.
The fifth position: Efficiency Heretics

The missing mechanism
The fifth position is defined more so by a claim about efficiency than about the source of knowledge, though, of course, the two are connected. It begins from an observation that a human child attains fluent command of a language from a quantity of exposure smaller by many orders of magnitude than a contemporary model requires, and that this disparity is the central unsolved problem: whatever a child possesses that permits it to learn so much from so little is precisely what current systems lack, and no increase in scale will solve that. The neolabs advancing this view frame it as a search for a missing mechanism rather than a missing dataset; one recently founded on exactly this premise, Flapping Airplanes, holds that human learners are more sample-efficient than current models by a factor somewhere between a hundred thousand and a million, and that closing that gap will require abandoning assumptions as deep as the mechanics of how these systems are trained.
This is (very clearly) Noam Chomsky's argument from the poverty of the stimulus, which was first advanced in the middle of the twentieth century, restated as an engineering program. Chomsky's case against statistical learning rests on the claim that the human mind "operates with small amounts of information," acquiring from impoverished data a high degree of competence. The neolabs in this camp have inherited the premise of Chomsky's argument (they accept that the sample-efficiency gap is the crux), but deny its conclusion (they hunt for a mechanism rather than accepting the idea of an innate grammar, which places them between the empiricists and the rationalists).
The standard objection to this premise is a good one. It states that the comparison of a child's exposure to a model's token count compares unlike things — the child's data is multimodal, interactive, socially scaffolded, and arrives in something like a curriculum, while the model's corpus is none of these. If that is where the child's advantage lives, the sample-efficiency gap is a data-quality gap rather than a mechanism gap, and the fifth position's central problem partially dissolves into the third position's program of richer, world-bound data. The camp's wager is that the disanalogy does not fully explain the gap; that wager may be right, but it is a wager, and I don't think its members always present it as one.
Adjacent to them are efforts to replace the sequential, one-token-at-a-time procedure by which most current models generate with alternatives that operate in parallel or through different dynamics. These are narrower architectural bets, but they share the intuition that the prevailing method is a contingent choice that can, and must, be challenged.
All told, this camp is unified by a diagnosis, making it the loosest of the five. Everyone here agrees the sample-efficiency gap is the real problem, sure, but no one agrees on what closes it. Some suspect the fix lies in the objective or the optimizer (Flapping Airplanes, for example, is willing to rethink both the loss function and gradient descent itself), while others suspect it lies in abandoning sequential, one-token-at-a-time generation altogether, whether through diffusion (Inception) or state-space models (Cartesia). Ultimately, I chose to place them by their shared complaint rather than by any common mechanism.
The instability of the positions
I should note that the positions I have mapped out here are fluid and should not be treated as a static taxonomy. Consider LeCun, whose vocabulary of architecture and inductive bias tempted me to file him among the rationalists of the fourth position. This classification is mistaken. LeCun is an empiricist about the content of knowledge (he holds that a system should learn everything from sensory data, positing no innate concepts) while wanting only a modest architectural scaffold to make that learning efficient. His position is a claim about form, in a mild and Kantian key, not the Fodorian claim about symbols with which it might be easily confused. As such, I think that he belongs in the third position, not the fourth. This confusion is so natural, though, that even Marcus has accused him of drifting toward a symbol-manipulation he once denounced.
Silver complicates the map from the opposite direction. He and LeCun are both adversaries of the first position, both convinced that text is insufficient, and it would be easy to enlist them together. But they divide on what is, in my view, the deepest axis of all — LeCun wants a measure of built-in architecture; Silver regards even that as an unearned concession, holding that the structure of the model itself ought to be discovered from experience.
The coalition against the language-first neolabs contains, in other words, a quarrel about how much structure is legitimate — and Silver stands at its far empiricist end. Friston, meanwhile, collapses the binary from within. Ilya Sutskever, who authored the first position's playbook, has by his own account recanted its strong form, migrating toward something that sounds increasingly like the second or fifth position; he is the hardest figure to place. And, as I noted above, the Scalers themselves have absorbed a large fraction of the Experientialists' methods while keeping the Scaler story, which is instability of a different kind and is happening in practice.
First consequence: the central dispute is a family quarrel
The dispute that occupies the field's attention and absorbs the overwhelming share of its capital is the confrontation between the first position and the third, or in other words, between those who would scale language models and those who would bind systems to the physical world. This is not, as I incorrectly assumed it would be, a contest between empiricism and rationalism at all. It is a disagreement between two species of empiricist, as both parties accept that intelligence is to be learned from data rather than furnished in advance, differing only over which data, text or world.
The second position is a third species of the same genus, disputing not whether to learn from experience but whose experience (humanity's or the system's own). So, as you can see, three of my five positions, and very nearly all of the money, fall on the empiricist side of the line. And the line is blurring even among them — the drift of the scaling labs toward experiential training means the family quarrel is being settled, in practice, by adopting the relative's methods and keeping one's own name.
The genuine rationalists are by contrast a periphery. But their marginality is perhaps better explained by the structure of the enterprise than by the state of their arguments (I am not in the position to do the latter). A position that predicts scale will fail is difficult to capitalize in a moment when scale is delivering very, very visible returns, and the neolabs that could be built around it are correspondingly few.
Second consequence: the theory of understanding that no one supplies
The second consequence is, in my view, the more significant, and it concerns something none of the five positions provides. Each offers an account of how intelligence is to be built (from text, from experience, from the world, from innate structure, from a more efficient mechanism) and each, when pressed on what the intelligence so built would understand, supplies a proxy. Understanding is operationalized as performance on a benchmark, success at a task, accuracy of prediction, or accumulated reward. This substitution is now well documented – the fact that a large share of the field's benchmarks purport to measure abstract capacities such as reasoning without defining what those capacities are is well established, and historians of the practice have traced how benchmarking came to displace theoretical argument as the field's mode of adjudication.
The obvious counterexample to this claim is interpretability, which is both well capitalized and aimed, apparently, at exactly the vacancy I am describing. But I don't think it fills it. Interpretability, as practiced, is a theory of mechanism. So it can tell you which circuits activate, which features a model represents, or what computation a given behavior runs on. What it cannot tell you (and frankly, what it does not even attempt to tell you) is what would make any such mechanism count as understanding. That question is prior to the reverse-engineering. A completed circuit-level description of a model would still leave open whether the model understands anything, for the same reason a completed wiring diagram of a brain would. So the field's most theory-shaped investment is an investigation of how the systems work, conducted in the continued absence of an account of what understanding is. (This is also why I read Anthropic's interpretability program as a hedge rather than a refutation. Think about it — it is the behavior of a lab that suspects prediction alone won't tell it what its models understand but not of one that has an answer.)
Why does the vacancy persist? I suspect that it is at least partly structural. A theory of understanding has no demo. You can capitalize a demonstration of what a system can do, but it is much harder to capitalize an argument about what doing it would mean. If that is right, the vacancy is not a lapse the field will eventually correct in the ordinary course of business, because the field's mode of raising money does not reward doing so. This is, I concede, an inference about incentives and not evidence about intentions, and I am sure that individual researchers at these labs care about the question a great deal. My claim is more limited to what the enterprise selects for (again, I am not in a position to make claims about what its members believe).
Two further observations sharpen my point here. The first is that the third position's central premise (that grounding requires the world, that a text-trained system is for that reason cut off from meaning) is the very claim that the most rigorous recent philosophy of grounding rejects. On the teleosemantic account, whether a representation is genuinely about the world turns on its standing in the right causal-informational relations and on a history of selection for carrying that information — conditions a text-trained system can in principle satisfy — and the authors' conclusion is that multimodality and embodiment are neither necessary nor sufficient for grounding. A quick caveat — teleosemantics is itself a contested program, with its own well-known difficulties, and one careful paper does not settle a debate. But perhaps it is telling that the best-funded challenge to the scaling neolabs rests on a premise that the sharpest available analysis of its own key concept treats as an open question.
My second observation is that the fifth position's sample-efficiency thesis is, as I have argued, Chomsky's argument in a new costume, and the migration of a sixty-year-old claim in linguistics into a funded engineering program, with a valuation attached to what was until recently a matter for seminar rooms, is a development the field has not paused to notice. Beneath both observations lies the same absence — a set of enormous wagers placed on the nature of understanding by parties none of whom has said what understanding is.
Limits of my analysis
The account I have given is an interpretation and there are of course several ways in which I could be wrong.
My most serious concern is that I have the adjudication backwards. I have written as though the burden lies on the language-first neolabs, but the empiricist case against them is strong and is arguably prevailing on the specific questions where it can be tested. Many have defended a moderate empiricism on which deep networks acquire structure from data without needing it installed in advance; Others have argued that modern language models undermine the innateness claims of generative linguistics outright; the work of some, such as Ellie Pavlick, keeps locating meaning-like internal structure where critics predicted only surface form; and, as I have already conceded, the strongest philosophy of grounding leans toward the scalers rather than against them.
It is entirely possible that the correct conclusion is not that the field misunderstands its disagreement, but that the empiricists are simply winning it, and that the rationalists are marginal because they are mistaken. My first consequence would survive this (the dispute would still be intramural to empiricism) but its tone of even-handedness toward the periphery would not.
My second concern is that this entire exercise imposes a philosophical vocabulary on what are, in the end, empirical bets. A defender of the field's self-understanding could say that a neolab choosing between text and video is not taking a position on the origins of the mind's contents but making a forecast about what will work, and that dressing the forecast in the language of Hume and Kant adds gravity but not the valuable content I was hoping it would. I don't think this objection succeeds, though, because the forecasts are underwritten by convictions about what kind of thing understanding is, which is something that no experiment has yet settled. But, I acknowledge that the line between a philosophical commitment and an empirical hypothesis is, in a young and fast-moving field, harder to draw than my argument has assumed.
My third concern is more local. The mapping I have done is contestable, and I have leaned on the instability of the positions as evidence while relying, to establish that instability, on readings of figures who are themselves in motion. A different reading of Sutskever, or of LeCun, or of what the second position's founders intend by grounded reward, would redraw parts of my map. And my claim about the revealed epistemology of the first position (that its practice is pragmatist whatever its rhetoric professes) is an inference from behavior to commitment that its subjects have not endorsed and might reasonably object to.
So my aim in all of this is the more modest one of showing that a set of disagreements the field treats as purely technical are perhaps structured by an older and deeper set of questions, and that seeing them as such reveals both a misdescribed alliance and a vacancy at the center of it all.
Of course, feedback and further discussion is always welcome :)
Further readings / What has been helpful for me
On neolabs
- The Neolab Wild West
- These Billion-Dollar AI Startups Have No Products, No Revenue and Eager Investors
From founders
- The Bitter Lesson
- Welcome to the Era of Experience
- AI and the Limits of Language
- From Words to Worlds: Spatial Intelligence is AI's Next Frontier
Philosophical debate
- Innateness and Language
- Large Language Models and the Rationalist-Empiricist Debate
- The False Promise of ChatGPT
- The Vector Grounding Problem
- From Deep Learning to Rational Machines
- Modern Language Models Refute Chomsky's Approach to Language
- Large Language Models and the Argument from the Poverty of the Stimulus
- Measuring What Matters: Construct Validity in Large Language Model Benchmarks
- AI as a Sport: On the Competitive Epistemologies of Benchmarking
