AGI Benchmarks: Eugenic Origins — by Valerie Veatch

Guest post by Valerie Veatch — some of her research notes for the film Ghost in the Machine.   The pursuit of ‘artificial general intelligence’ (AGI), or more recently ‘super intelligence’, has dominated research and development investment in recent years in Machine Learning (ML). “Achieving…

Guest post by Valerie Veatch — some of her research notes for the film Ghost in the Machine.   The pursuit of ‘artificial general intelligence’ (AGI), or more recently ‘super intelligence’, has dominated research and development investment in recent years in Machine Learning (ML). “Achieving human-level ‘intelligence’ is an implicit or explicit north-star goal for many”[1] of the major Silicon Valley tech companies such as Deepmind (founded in 2010, now Google Deepmind)[2], OpenAI (founded in 2015 with the aim to “feel the AGI”), Anthropic (born out of OpenAI in 2021) and Meta (formerly Facebook, launching their product publicly in 2023). Though terminology varies, these companies have all harnessed their efforts around winning what is colloquially referred to as the ‘race to build AGI’, human level intelligence, or ‘super-intelligence’, receiving billions in funding from the private sector and governments, and driving massive global investment in data infrastructure. Understanding the origins of the terms ‘AGI’ as well as ‘super intelligence’, how these terms are defined, and the metrics around how the achievement of ‘AGI’ or ‘super intelligence’ are evaluated provides a sobering view into the legacy of measuring and defining intelligence in humans, which has always been a political act. How such intelligence is defined, measured, and ultimately determined carry huge stakes for billions — even trillions[3] — in funding as well as the public perception of the functionality of ‘AI’ systems and products as they are relentlessly deployed. By declaring ‘AGI’ or human-level intelligence AI developers shift the narrative around policy and public perception of AI capabilities and risks. Modern benchmarks for ‘AGI’ such as ARC-AGI[4], AGIEval, MMLU, and BIG-bench — aim to quantify an AI system’s ‘general’ intelligence by testing it on a broad range of tasks. These benchmarks echo the structure and philosophy of human IQ and standardized tests, treating intelligence as a measurable, comparable quantity. This is not coincidental: many AGI evaluations explicitly draw on psychometric methods (puzzles, exams, test batteries) that were originally developed to measure human intelligence for the purposes of eugenics. Crucially, those human intelligence tests arose in the early 20th century under the influence of hereditarian and eugenic theories — the belief that intelligence is a single innate property of individuals (and races) that can be quantified and ranked[5]. Pioneers of psychometrics like Francis Galton, Charles Spearman, Lewis Terman, and Carl Brigham were active in the eugenics movement, using tests to promote ideas of racial hierarchy and genetic merit. The legacy of that era still permeates how we define and measure the intelligence of machines today[6]. Defining AGI: intelligence and white nationalism As we will see in subsequent sections of the paper, the term ‘general intelligence’ is soaked in 20th century eugenics and white nationalism[7]. Here we will focus on the recent efforts to apply the term to ML and the numerous uses of white nationalist and eugenicist scholars to define ‘intelligence’. The notion of a ‘general’ intelligence in relation to ML typically refers to systems that can supplant human intelligence. Artificial general intelligence was invoked by Mark Gubard describing a system that would equate or exceed the capacities of a human soldier in a battlefield context in November 1997 at the Nano Technology and Security Conference[8]: “By advanced artificial general intelligence, I mean AI systems that rival or surpass the human brain in complexity and speed, that can acquire, manipulate and reason with general knowledge, and that are usable in essentially any phase of industrial or military operations where a human intelligence would otherwise be needed.” In the context of the current race to build ‘AGI’ the phrase ‘Artificial General Intelligence’ was independently coined and popularised in 2001 by Shane Legg (DeepMind) and Ben Goertzel. As Legg describes[9]; he was approached by Goertzel about a book on the theme of human-like intelligence, Legg suggested the title Artificial General Intelligence. The opening chapter of this volume, written by Goertzel, seeks to define “AGI” and cites a book The g Factor: General Intelligence and Its Implications, written by Christopher Brand (at the time was a fellow at the Galton Institute known for making controversial statements linking poverty to hereditary race) published and de-published in 1996 due to public outcry around the blatantly racist elements of the book[10]. This was the beginning of the relentless use of white-nationalist, white-supremecist, eugenicist definitions of ‘intelligence’ deployed by the computer sciences to define what eventual machine ‘general intelligence’ might look like. To this day Legg remains at the frontier in defining and shaping what ‘AGI’ might mean. We see in his 2008 doctoral dissertation the way he frames “such questions of ‘are whites smarter than blacks, are men smarter than women’” as “controversial” and it’s best to steer away from addressing such nuances when exploring how to define intelligence in the service of the project to create super intelligent machines. For Legg this acknowledgement seemingly unburdens him from repeated deployment of racist definitions of intelligence. He continuously moves ever into the center of the mainstream and utilises thinking and citations derived directly from eugenics and white nationalism to perpetuate this framing of intelligence as it’s applied to large algorithmic systems. Legg and Marcus Hutter (Deepmind) in a December 2007 list of 70 definitions of intelligence[11] reference several definitions of intelligence conflating intelligence with race. This list of 70 definitions included a definition by David Wecheister, who during WW1 was drafted to work under famed eugenists Charles Spearman and Karl Pearson to develop psychological screening tests for new draftees organised by race which later led to the Immigration Act of 1924, which included quotas for immigration to the United States[12] (see: fig 5 in Section 2.6), in his 1959 book Appraisal and Measurement of Adult Intelligence. Another definition invoked by Legg numerous times in various papers was by white nationalist[13] Linda Gottfredson, in an op-ed for the Wall Street Journal in defense of “The Bell Curve” titled “Mainstream science on intelligence” in which Gottfredson asserts “blacks have a lower IQ than whites for[…] hereditary reasons.”[14] This definition is framed by Legg and Hunter as “a definition of intelligence agreed upon by 53 psychologists” some of whom were funded by the Pioneer Fund, along with Gottfredson, a neo Nazi eugenicist white supremacist organisation, classified as a hate group by the Southern Poverty Law Center, which also describes Gottfredson as a white nationalist[15]. In Legg’s doctoral thesis published the following year for the University of Lugon, Legg cited Gottfredson thrice in his thesis including her 1997 piece “Mainstream science on Intelligence”, a piece called “Why G Matters”[16][17]. Those definitions of “general” intelligence overtly ascribed a hereditary causal relationship to ‘intelligence’. This doctoral thesis was overseen by Hutter (currently Deepmind, formerly University of Lugon). Despite the rigours of a thesis review process somehow citations of multiple works by a prominent white nationalist remained, either from ignorance or lack of care. In 2023 this definition by Gottfredson again appeared as the only definition for intelligence in a 2023 Microsoft paper entitled “Sparks of AGI: experiments with Chat GPT” (versions 1,2,3,4[18]). The paper was later updated five times, in the final version removing the direct reference to Gottfredson, replaced with: “There is no generally agreed upon definition of intelligence, but one aspect that is broadly accepted is that intelligence is not limited to a specific domain or task, but rather encompasses a broad range of cognitive skills and abilities”[19]. In a 2024 Google Deepmind paper “Levels of AGI” (co-authored by Legg) the authors hail the use of benchmarks to track progress, and cite the Microsoft paper “Sparks of AGI” as reference to the impending declaration of ‘AGI’. It should be noted that even as companies shift their language around “AGI” or “super-intelligence” the goal remains steadfast: create intelligence that reaches or exceeds human intelligence. Before we continue it’s useful to examine how each of these major players define AGI today and how they frame the achievement of AGI. OpenAI Define AGI: OpenAI’s charter frames AGI in capitalist terms as “highly autonomous systems that outperform humans at most economically valuable work.”[20] Measure AGI: Benchmarks such as ARC-AGI test, performance parity with skilled humans, external audits, and governmental oversight. Intend to Declare AGI: When systems consistently outperform humans across major benchmarks, receive external verification, and clear safety assessments. DeepMind Define AGI: Levels of AGI framework (Emerging, Competent, Expert, Virtuoso, Superhuman) based on task breadth and depth. Measure AGI: Rigorous ecological validity tests and continuous measurement across diverse cognitive tasks, risk assessment at each capability level. Intend to Declare AGI: Avoids declaring a single AGI moment; instead, gradually acknowledges progression through clearly defined capability levels via benchmarks Anthropic Define AGI: Avoids AGI term; prefers “powerful AI,” defined as superior general reasoning and capability across economic and scientific tasks. Measure AGI: AI Progress Index focusing on economic and scientific task performance, robust general reasoning and knowledge across tasks. Intend to Declare AGI: Unlikely to declare explicit AGI; may acknowledge powerful AI once models exceed human performance across broad, critical fields and benchmarks around 2026. Meta Define AGI: In 2024 Meta frames developing ‘AGI’ as a long term goal, as of March 2025 prefers the term ‘advanced machine intelligence.’[21][22] Measure AGI: Emphasizes real-world problem-solving abilities, human-level performance in commonsense reasoning, and tasks requiring embodiment. Intend to Declare AGI: Meta’s chief scientist images ‘AGI’ will be available in “3-5 years” from 2025 as measured by improvements towards human-level capabilities in specific domains via benchmarking. Evaluation ARC-AGI[23], AGIEval, MMLU, and BIG-bench In addition to the specific tests within each benchmark, there are broader assumptions that linger from eugenics in the endeavor to declare “AGI’. The reification of intelligence as a single metric or a small set of metrics (e.g., “GPT-4 has a 90% on MMLU, so it’s very intelligent” parallels “Person A has IQ 130, so they’re very intelligent”) relies on the belief in test-based objectivity, that complex qualities of cognition can be fairly captured by well-designed questions and scored impartially. The use of these metrics to guide high-stakes decisions have a sordid history: In humans it was school placement, immigration, reproduction rights (tragically, Buck v. Bell 1927 legitimized sterilizing a woman labeled “feeble-minded” largely on such test-based judgment). [24] In AI, the stakes are different but significant: definitions of intelligence have been weaponized to oppress, from colonial times where indigenous peoples were deemed “inferior intellects,” to the 20th-century scientific racism that IQ tests fueled. [25] If our AI benchmarks inherit those same definitions, we might inadvertently build systems that excel at a very narrow, Western, upper-class notion of intelligence — and deem other forms of reasoning as less important, repeating the marginalization in a new domain. If an AI is deemed super-intelligent by these tests, will it be trusted too much with decisions that affect people? If the benchmarks don’t include ethical reasoning, a model could be very “smart” by score but make harmful choices. In the analysis below, we examine each major ‘AGI’ benchmark in detail, how it defines ‘AGI’, the design of its tasks, and how its evaluation criteria trace back to traditional IQ tests and standardized exams. We map each benchmark onto its psychometric lineage, identifying the historical tests and figures that inspired it and highlighting the ideologies embedded in those tools. We also incorporate scholarly critiques (from cognitive science, AI ethics, and critical race theory) to question the assumptions and biases that may be carrying forward. For simplicity, the following maps each AGI benchmark to its closest psychometric analog and eugenic-era origin: ARC-AGI (Francois Chollet, 2019, Google) AGI Focus & Task Design: Defines intelligence as “general fluid intelligence” or skill-acquisition ability. Consists of novel visual puzzles (grid pattern tasks) that must be solved from only a few examples.[26] Measures an AI’s ability to generalize from minimal data (like a human solving an IQ puzzle cold). Psychometric Lineage & Eugenic Origins: Raven’s Progressive Matrices (non-verbal IQ test, 1938) is the direct inspiration.[27] ARC’s colored grid puzzles mirror Raven’s pattern matrices, which were designed to measure abstract reasoning (fluid intelligence) independent of language. Chollet’s documentation states “the task format is inspired by Raven’s Progressive Matrices”. Raven’s test was developed under Charles Spearman’s theory of a unitary g factor of intelligence and was touted as a “culture-fair” way to test innate reasoning. Spearman and contemporaries (e.g. Francis Galton and Karl Pearson) were prominent eugenicists; their statistical tools (like correlation) and belief in an inherited, immutable intelligence shaped the very format of such tests. Raven’s and similar puzzles were used by militaries and psychologists to identify “feeble-mindedness” vs. genius, aligning with eugenic goals of stratifying people by presumed genetic intellect. Chollet’s ARC explicitly positions itself as a “psychometric intelligence test” for AI , inheriting that lineage of testing abstract reasoning as the core of “general” intelligence. AGIEval (Wanjun Zhong et al., 2023, Microsoft Research) AGI Focus & Task Design: Frames AGI in terms of performance on human academic and professional exams. Uses real standardized exams (e.g. college entrance tests, the SAT, law school LSAT, math Olympiad, China’s college Gaokao, bar exams) to evaluate AI[28]. Success = scoring as well as or better than human test-takers on these knowledge and reasoning tests. Psychometric Lineage & Eugenic Origins: Standardized aptitude tests in the 20th century (e.g. SAT, LSAT, civil service exams) form the template. The SAT (Scholastic Aptitude Test) was created in 1926 by Carl Brigham, who had worked on the U.S. Army’s WWI IQ tests and was an open eugenicist. Brigham believed such tests could rank innate ability across racial/ethnic groups and initially used Army test data to argue for Nordic racial superiority in intelligence. The LSAT (1948) and other entrance exams followed the SAT’s approach: multiple-choice questions testing logic, verbal reasoning, and math under time pressure, all methods rooted in psychometric research. Lewis Terman (who popularized the Stanford-Binet IQ test in 1916) had explicitly called for using IQ-style tests to sort individuals into educational and professional tracks, even suggesting minimum IQ cutoffs for various occupations. He and fellow eugenicists (like Robert Yerkes and Henry Goddard) pushed standardized testing into schools, the military, and immigration, arguing that low scores indicated hereditary “feeble-mindedness” deserving of social exclusion or sterilization. AGIEval directly uses these tests as benchmarks for AI, implicitly adopting the premise that performance on exams equates to intelligence. The very format and content of exams like the SAT and LSAT, (analogies, reading comprehension, logic puzzles), descend from early IQ test components, many of which were created by eugenicists to prove mental differences. Brigham’s racist assumptions and Terman’s hereditarian views are part of the DNA of these exams, and thus they persist as silent backdrops when we treat exam results as a measure of “general” intelligence. MMLU (Dan Hendrycks et al., 2021, UC Berkeley/OpenAI) AGI Focus & Task Design: Tests an AI’s breadth of knowledge and problem-solving across 57 diverse subjects. Includes elementary math, US history, computer science, law, chemistry, biology, etc. Each subject is a set of factual and analytical multiple-choice questions (modeled after school and professional quizzes). An AGI-level model would be expected to score highly in all topics, demonstrating broad academic-level competence[29]. Psychometric Lineage & Eugenic Origins: Comprehensive IQ batteries and achievement tests inform this benchmark. By covering both academic knowledge and analytical reasoning, MMLU spans what psychologists call crystallized vs. fluid intelligence. This distinction was introduced by Raymond Cattell (a student of Spearman and himself a proponent of eugenic ideas), where fluid intelligence (Gf) is solving novel problems (tested by puzzles like Raven’s), and crystallized intelligence (Gc) is accumulated knowledge (tested by vocabulary, factual tests, etc.). MMLU’s trivia and problem sets parallel an amalgam of an IQ test (for reasoning) and a college board exam (for learned knowledge). Historically, after IQ testing took hold, the same multiple-choice, standardized format was applied to curricular knowledge: e.g. the Iowa Tests of Basic Skills and the expansion of the SAT into subject tests. The creators of those tests, often working at institutions like ETS (Educational Testing Service) founded by Brigham and colleagues, believed in objective ranking of students similar to IQ rankings. Terman had envisioned a society where, by testing everyone’s mental abilities, individuals could be efficiently slotted into roles, an idea that dovetailed with capitalist and eugenic notions of “merit”. MMLU’s design, a broad battery graded by percentage correct, inherits the assumption that a single aggregated score across domains is a valid indicator of general intelligence or academic ability. This is exactly how IQ tests (like Wechsler’s scales) combine sub-test scores into one IQ number. Importantly, the content of MMLU (mostly English-language, Eurocentric academic knowledge) reflects the implicit bias of which knowledge counts as a marker of intelligence — a bias rooted in the Euro-American educational canon established by those early test-makers (who often excluded or devalued non-Western knowledge systems). The eugenic context of early achievement testing, for example, using test performance to justify excluding immigrants and minorities from higher education, raises caution that high scores may signal alignment with a very specific, inherited notion of “intelligence,” rather than neutral, universal genius. BIG-bench (Beyond the Imitation Game, 2022, multi-institution collaboration led by Google) AGI Focus & Task Design: Proposes a diverse collection of 204 tasks to evaluate AI, going beyond a single test. Tasks range from language understanding (translation, question-answering) to reasoning puzzles (logic grid problems, mathematical word problems) to creative tasks (jokes, riddles). The benchmark’s goal is to “probe and extrapolate” a model’s general capabilities across many domains, under the premise that an AGI should handle any task that a human intellect could. Models are scored on each task and compared to human performance, to identify areas of strength and weakness. Psychometric Lineage & Eugenic Origins: IQ test batteries & the “g factor” theory are the clear precursors. BIG-bench is essentially a battery of tests for AI, analogous to how human intelligence research uses a battery of sub-tests to measure different cognitive abilities. The underlying rationale is very much Spearman’s: if a system truly has high general intelligence, it should perform well across all kinds of tasks (not just one skill). In psychometrics, this idea was operationalized by combining varied sub-tests (verbal analogies, arithmetic, spatial puzzles, etc.) into a single evaluation, and statistically extracting a dominant factor (g) from the correlations. Spearman developed this approach in the early 1900s amidst a fervent hope that intelligence testing could scientifically prove innate hierarchies (he described g as the “mental energy” distinguishing individuals, often suggesting it was hereditary). BIG-bench’s diverse-task format mirrors tests like Wechsler’s Adult Intelligence Scale, where performance on diverse tasks is aggregated. Moreover, many BIG-bench tasks directly resemble classic IQ or aptitude test items: for example, it includes logical deduction problems and analogies very similar to those on the SAT/GRE and the older Army Alpha tests. Those original test items were created by psychologists like Arthur Otis (developer of Army Alpha’s analogies) and Brigham, explicitly to rank individuals by cognitive ability for the Army and universities — endeavors entangled with eugenic sorting (Army testing data was famously misused to claim racial mental deficits). Ideologically, BIG-bench perpetuates the notion that intelligence equals proficiency on a sum of academic and puzzle challenges. This notion has been challenged by cognitive scientists who argue intelligence is multifaceted and context-dependent (e.g. Howard Gardner’s theory of multiple intelligences, which notes that skills like interpersonal or musical intelligence aren’t captured by pen-and-paper tests). Yet, BIG-bench (by design) omits many real-world intelligences (physical, social, emotional), focusing on the types of “academic” intelligence that early psychometricians (often white, Western men) considered supreme. In that sense, it carries forward a partial and culturally specific picture of intelligence. The very name “Beyond the Imitation Game” signals a move past the Turing Test towards quantitative evaluation echoing how 20th-century psychologists moved from subjective assessments of intelligence to formal tests. Contemporary AGI benchmarks are deeply rooted in the paradigms of psychometric testing. ARC-AGI channels the spirit of Spearman and Raven, focusing on abstract reasoning puzzles to capture “fluid intelligence” in machines. AGIEval turns Brigham’s SAT and other standardized exams (born of the eugenics-infused early 20th century) into hurdles for AI, implicitly endorsing those tests’ claims to measure aptitude. MMLU compiles a tapestry of academic subjects, reflecting the longstanding practice of gauging intellect by the breadth of one’s learned knowledge — an approach with origins in the achievement tests and trivia contests that themselves are aligned with IQ assumptions about who can acquire knowledge. BIG-bench assembles an extensive test battery for AI, a direct parallel to an IQ test battery, emanating from the belief that general intelligence manifests in performance across many tasks. The ideological lineage of these benchmarks traces back to figures like Francis Galton, Charles Spearman, Alfred Binet, Lewis Terman, Raymond Cattell, and Carl Brigham. These individuals laid the foundations of quantifying intelligence, often with the explicit aim of proving hierarchical differences between people. Galton’s eugenic vision and statistical innovations set the stage for viewing intelligence as a heritable, rankable trait. Spearman’s g factor gave scientific cover to the idea of a single scale of “smartness,” which eugenicists eagerly applied to justify stratified social policies. Terman and Brigham developed and deployed tests (Stanford-Binet, Army Alpha, SAT) that were used to elevate some and marginalize others, firmly convinced (at least early on, in Brigham’s case) that these scores reflected genetic destiny. Cattell’s introduction of fluid vs crystallized IQ lives on in benchmarks like ARC (fluid) versus MMLU (crystallized),and Cattell too entwined his scientific work with eugenic philosophy later in life.. Modern AGI benchmarks can be seen as the progeny of IQ and standardized tests: they carry forward the structural DNA and many underlying values of their ancestors. This isn’t to say they are useless or ill-intentioned, they have driven rapid advances and identified weaknesses in AI models. However, recognizing their psychometric and eugenic lineage is crucial to understanding their limitations and how we frame the hype around AGI. As Dixon-Román notes, our “sociotechnical systems of quantification” have always been entwined with power and bias. By tracing ARC, AGIEval, MMLU, and BIG-bench back to their roots in tests designed by eugenicists, we become better equipped to question what these benchmarks are really measuring. Are we inadvertently baking in the same old hierarchies and blind spots found in eugenic testing whilst enabling a ludicrous concept that a machinic algorithmic system might equate to human intelligence?     Note and source: Above, OpenAI celebrates the launch of another generative image product, “FEEL THE AGI” text appears over a Studio Ghibli-style generated image, middle figure flashing the “OK” or white power hand symbol. Below, a comment on the post noting the image generator “can’t even white power hand sign right.” (x.com/sama) 26 March 2025.   1 https://ar5iv.labs.arxiv.org/html/2311.02462#:~:text=6,example%2C%20Aguera%20y%20Arcas 2 https://deepmind.google/about/ 3 https://www.ft.com/content/1fda45a2-43e0-4c10-b5fb-b6097e3f5c56 4 https://github.com/fchollet/ARC-AGI 5 Graves, J. L., & Johnson, A. (1995). The Pseudoscience of Psychometry and The Bell Curve. The Journal of Negro Education, 64(3), 277—294. https://doi.org/10.2307/2967209 6 Wintroub, M. Sordid genealogies: a conjectural history of Cambridge Analytica’s eugenic roots. Humanit Soc Sci Commun 7, 41 (2020). https://doi.org/10.1057/s41599-020-0505-5 7 Yarden Katz, 2021 “Intelligence Under Racial Capitalism from Eugenics to Standardized Testing and Online Learning” https://monthlyreview.org/2022/09/01/intelligence-under-racial-capitalism-from-eugenics-to-standardized-testing-and-online-learning/#:~:text=The%20quest%20to%20pin%20down,%E2%80%9D 8 https://legacy.foresight.org/Conferences/MNT05/Papers/Gubrud/index.html 9 “The Transformative Potential of AGI — and When It Might Arrive | Shane Legg and Chris Anderson | TED”, @01:49 https://www.youtube.com/watch?v=kMUdrUP-QCs 10 https://en.wikipedia.org/wiki/The_g_Factor:_General_Intelligence_and_Its_Implications 11 https://arxiv.org/pdf/0706.3639 12 https://en.wikipedia.org/wiki/Immigration_Act_of_1924 13 https://www.splcenter.org/resources/extremist-files/linda-gottfredson/ 14 Gottfredson “Mainstream Science on Intelligence” https://psycnet.apa.org/record/1997-38920-001 15 https://www.splcenter.org/resources/extremist-files/linda-gottfredson/ 16 Gottfredson “Why G Matters” https://www1.udel.edu/educ/gottfredson/reprints/1997whygmatters.pdf 17 http://www.vetta.org/documents/Machine_Super_Intelligence.pdf (retrieved 27 March 2025) 18 “Sparks of AGI” Version 1 remains on arXiv, Version 5 has replaced the Linda Gottfredson definition of intelligence with this sentence: “There is no generally agreed upon definition of intelligence, but one aspect that is broadly accepted is that intelligence is not limited to a specific domain or task, but rather encompasses a broad range of cognitive skills and abilities.” 19 “Sparks of AGI: Experiments with Chat GPT” Version 5, 13 April 2023 https://arxiv.org/pdf/2303.12712 20 https://openai.com/charter/ 21 https://www.nationaltechnology.co.uk/Meta_Chief_AI_Scientist_Claims_AGI_Will_Be_Viable_In_3_5_Years.php 22 https://www.axios.com/2024/01/22/meta-artificial-general-intelligence-quest-agi 23 https://github.com/fchollet/ARC-AGI 24 Buck v. Bell, 274 U.S. 200 (1927) 25 Katz, Yardin https://monthlyreview.org/2022/09/01/intelligence-under-racial-capitalism-from-eugenics-to-standardized-testing-and-online-learning/#:~:text=The%20quest%20to%20pin%20down,%E2%80%9D 26 Chollet, F. (2019). On the Measure of Intelligence — Introduces ARC as a psychometric test for AI, discusses definitions of intelligence. Retrieved 17 April 2025 27 Freethink (2023). “LLMs are a dead end to AGI, says François Chollet” — Article where Chollet compares ARC to “a human IQ test invented in 1938, called Raven’s Progressive Matrices.” Includes an example ARC puzzle and Chollet’s critiques of LLMs. Retrieved 17 April 2024 28 Zhong, W. et al. (2023). AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models — Describes the use of standardized exams (SAT, LSAT, etc.) to evaluate AI (GPT-4, ChatGPT) performance. Retrieved 17 April 2025 29 Hendrycks, D. et al. (2021). Measuring Massive Multitask Language Understanding — Presents MMLU with 57 subject areas to assess broad knowledge and reasoning in models; notes model performance and shortcomings.

Source: Pivot to AI — Published — Category: Business

🔗 Read full article on Pivot to AI →