Search AI News
Find articles from 100+ AI sources
Found 298 results for "LLaMA"
State of AI: December 2025 newsletter
Dear readers, Welcome to the latest issue of the State of AI, an editorialized newsletter that covers the key…
Beyond Standard LLMs
From DeepSeek R1 to MiniMax-M2, the largest and most capable open-weight LLMs today remain autoregressive…
Introducing Gemma 3n: The developer guide
The first Gemma model launched early last year and has since grown into a thriving Gemmaverse of over 160…
🪩 The State of AI Report 2025 🪩
Hi everyone!The day is finally here: I’m thrilled to share the State of AI Report 2025 with you!In short,…
Understanding the 4 Main Approaches to LLM Evaluation (From Scratch)
How do we actually evaluate LLMs?It’s a simple question, but one that tends to open up a much bigger…
REINFORCE: Easy Online RL for LLMs
Reinforcement learning (RL) is playing an increasingly important role in research on large language models…
Online versus Offline RL for LLMs
(from [2, 5, 7, 9, 10])The alignment process teaches large language models (LLMs) how to generate completions…
H100 vs GB200 NVL72 Training Benchmarks – Power, TCO, and Reliability Analysis, Software Improvement Over Time
Frontier model training has pushed GPUs and AI systems to their absolute limits, making cost, efficiency,…
From GPT-2 to gpt-oss: Analyzing the Architectural Advances
OpenAI just released their new open-weight LLMs this week: gpt-oss-120b and gpt-oss-20b, their first…
Direct Preference Optimization (DPO)
(from [1, 2, 6, 9])Aligning large language models (LLMs) is a crucial post-training step that ensures models…
Migrating the Hub from Git LFS to Xet
Back to Articles Migrating the Hub from Git LFS to Xet Published July 15, 2025 Update on GitHub Upvote 29 +23…
State of AI: July 2025 newsletter
Hi everyone!Welcome to the latest issue of the State of AI newsletter, an editorialized newsletter covering…
LinkedIn Highlights, June 2025 - AI Agents Edition
Welcome to LinkedIn Highlights!Each month, I'll share my five seven top-performing LinkedIn posts, bringing…
Meta Superintelligence – Leadership Compute, Talent, and Data
Meta’s shocking purchase of 49% of Scale AI at a ~$30B valuation shows that money is of no concern for the…
LLM Research Papers: The 2025 List (January to June)
As some of you know, I keep a running list of research papers I (want to) read and reference.About six months…
Reward Models
(from [1, 2, 4, 14])Reward models (RMs) are a cornerstone of large language model (LLM) research, enabling…
Transformers backend integration in SGLang
Back to Articles Transformers backend integration in SGLang Published June 23, 2025 Update on GitHub Upvote…
Understanding and Coding the KV Cache in LLMs from Scratch
KV caches are one of the most critical techniques for efficient inference in LLMs in production. KV caches…
MIT Researchers Unveil “SEAL”: A New Step Towards Self-Improving AI
The concept of AI self-improvement has been a hot topic in recent research circles, with a flurry of papers…
How Long Prompts Block Other Requests - Optimizing LLM Performance
Back to Articles How Long Prompts Block Other Requests - Optimizing LLM Performance Team Article Published…
State of AI: June 2025 newsletter
Hi everyone!Welcome to the latest issue of the State of AI newsletter, an editorialized newsletter covering…
Dell Enterprise Hub is all you need to build AI on premises
Back to Articles Dell Enterprise Hub is all you need to build AI on premises Published May 23, 2025 Update on…
A Guide for Debugging LLM Training Data
(from [2])Most discussions of LLM training focus heavily on models and algorithms. We enjoy experimenting…
DeepSeek-V3 New Paper is coming! Unveiling the Secrets of Low-Cost Large Model Training through Hardware-Aware Co-design
A newly released 14-page technical paper from the team behind DeepSeek-V3, with DeepSeek CEO Wenfeng Liang as…