Search AI News
Find articles from 100+ AI sources
Found 1,049 results for "Google"
Kimina-Prover: Applying Test-time RL Search on Large Formal Reasoning Models
Back to Articles Kimina-Prover: Applying Test-time RL Search on Large Formal Reasoning Models Team Article…
Against "Brain Damage"
I increasingly find people asking me “does AI damage your brain?” It's a revealing question. Not because…
Announcing NeurIPS 2025 E2LM Competition: Early Training Evaluation of Language Models
Back to Articles Announcing NeurIPS 2025 E2LM Competition: Early Training Evaluation of Language Models…
Reward Models
(from [1, 2, 4, 14])Reward models (RMs) are a cornerstone of large language model (LLM) research, enabling…
Researchers from PSU and Duke introduce “Multi-Agent Systems Automated Failure Attribution
Share My Research is Synced’s column that welcomes scholars to share their own research breakthroughs with…
AI metrics
In the early days of the consumer Internet, a lot of metrics floated around and no-one was clear what to…
AI Agents from First Principles
(from [1] and source)The capabilities of large language models (LLMs) are advancing rapidly. As LLMs become…
State of AI: June 2025 newsletter
Hi everyone!Welcome to the latest issue of the State of AI newsletter, an editorialized newsletter covering…
GenAI’s adoption puzzle
This chart is very ‘glass half-empty or half-full?’, and it’s a puzzle. You could say that this is…
Dell Enterprise Hub is all you need to build AI on premises
Back to Articles Dell Enterprise Hub is all you need to build AI on premises Published May 23, 2025 Update on…
A Guide for Debugging LLM Training Data
(from [2])Most discussions of LLM training focus heavily on models and algorithms. We enjoy experimenting…
State of AI: May 2025 newsletter
Hi everyone!Welcome to the latest issue of the State of AI newsletter, an editorialized newsletter covering…
LinkedIn Highlights, Apr 2025
Welcome to LinkedIn Highlights!Each month, I'll share my five top-performing LinkedIn posts, bringing you the…
PGN2FEN: A Benchmark for Evaluating LLM Chess Reasoning
Today, I’m releasing PGN2FEN — a new benchmark that tests the ability of language models to understand…
Introducing HELMET: Holistically Evaluating Long-context Language Models
Back to Articles Introducing HELMET: Holistically Evaluating Long-context Language Models Published April 16,…
Full speaker lineup for RAAIS 2025
The Research and Applied AI Summit (RAAIS) is a community for entrepreneurs and researchers advancing the…
Hugging Face to sell open-source robots thanks to Pollen Robotics acquisition 🤖
Back to Articles Hugging Face to sell open-source robots thanks to Pollen Robotics acquisition 🤖 Published…
Visual Salamandra: Pushing the Boundaries of Multimodal Understanding
Back to Articles Visual Salamandra: Pushing the Boundaries of Multimodal Understanding Team Article Published…
Defending against Prompt Injection with Structured Queries (StruQ) and Preference Optimization (SecAlign)
Recent advances in Large Language Models (LLMs) enable exciting LLM-integrated applications. However, as LLMs…
Arabic Leaderboards: Introducing Arabic Instruction Following, Updating AraGen, and More
Back to Articles Arabic Leaderboards: Introducing Arabic Instruction Following, Updating AraGen, and More…
State of AI: April 2025 newsletter
Hi everyone!Welcome to the latest issue of the State of AI newsletter, an editorialized newsletter covering…
When machines learn to speak
This post is part of my 2¢ series - my raw thoughts about recent topics in AI. Not always practical…
Staking your ground
Hi everyone,As the AI world navigates through a choppy public market backdrop, I wanted to share an essay I…
How To Construct Self-Attention Mechanisms For Arbitrary Long Sequences
With Gemini models having a 2M tokens context size and Claude having a 200K tokens context size while having…