⚡ LIVE

Search AI News

Find articles from 100+ AI sources

🔍

Found 298 results for "LLaMA"

298 articles
Air Street Capital Business Nov 30, 2025 👁 43

State of AI: December 2025 newsletter

Dear readers, Welcome to the latest issue of the State of AI, an editorialized newsletter that covers the key…

AI Magazine (Raschka) Models Nov 4, 2025 👁 51

Beyond Standard LLMs

From DeepSeek R1 to MiniMax-M2, the largest and most capable open-weight LLMs today remain autoregressive…

DeepMind Research Oct 25, 2025 👁 41

Introducing Gemma 3n: The developer guide

The first Gemma model launched early last year and has since grown into a thriving Gemmaverse of over 160…

Air Street Capital Business Oct 9, 2025 👁 59

🪩 The State of AI Report 2025 🪩

Hi everyone!The day is finally here: I’m thrilled to share the State of AI Report 2025 with you!In short,…

AI Magazine (Raschka) Models Oct 5, 2025 👁 46

Understanding the 4 Main Approaches to LLM Evaluation (From Scratch)

How do we actually evaluate LLMs?It’s a simple question, but one that tends to open up a much bigger…

Cameron Wolfe (AI) Models Sep 29, 2025 👁 49

REINFORCE: Easy Online RL for LLMs

Reinforcement learning (RL) is playing an increasingly important role in research on large language models…

Cameron Wolfe (AI) Models Sep 8, 2025 👁 54

Online versus Offline RL for LLMs

(from [2, 5, 7, 9, 10])The alignment process teaches large language models (LLMs) how to generate completions…

SemiAnalysis Models Aug 20, 2025 👁 58

H100 vs GB200 NVL72 Training Benchmarks – Power, TCO, and Reliability Analysis, Software Improvement Over Time

Frontier model training has pushed GPUs and AI systems to their absolute limits, making cost, efficiency,…

AI Magazine (Raschka) Models Aug 9, 2025 👁 44

From GPT-2 to gpt-oss: Analyzing the Architectural Advances

OpenAI just released their new open-weight LLMs this week: gpt-oss-120b and gpt-oss-20b, their first…

Cameron Wolfe (AI) Models Jul 28, 2025 👁 43

Direct Preference Optimization (DPO)

(from [1, 2, 6, 9])Aligning large language models (LLMs) is a crucial post-training step that ensures models…

Hugging Face Open Source Jul 15, 2025 👁 42

Migrating the Hub from Git LFS to Xet

Back to Articles Migrating the Hub from Git LFS to Xet Published July 15, 2025 Update on GitHub Upvote 29 +23…

Air Street Capital Business Jul 13, 2025 👁 44

State of AI: July 2025 newsletter

Hi everyone!Welcome to the latest issue of the State of AI newsletter, an editorialized newsletter covering…

AI Tidbits Business Jul 13, 2025 👁 45

LinkedIn Highlights, June 2025 - AI Agents Edition

Welcome to LinkedIn Highlights!Each month, I'll share my five seven top-performing LinkedIn posts, bringing…

SemiAnalysis Models Jul 11, 2025 👁 42

Meta Superintelligence – Leadership Compute, Talent, and Data

Meta’s shocking purchase of 49% of Scale AI at a ~$30B valuation shows that money is of no concern for the…

AI Magazine (Raschka) Models Jul 1, 2025 👁 47

LLM Research Papers: The 2025 List (January to June)

As some of you know, I keep a running list of research papers I (want to) read and reference.About six months…

Cameron Wolfe AI Models Jun 30, 2025 👁 48

Reward Models

(from [1, 2, 4, 14])Reward models (RMs) are a cornerstone of large language model (LLM) research, enabling…

Hugging Face Open Source Jun 23, 2025 👁 40

Transformers backend integration in SGLang

Back to Articles Transformers backend integration in SGLang Published June 23, 2025 Update on GitHub Upvote…

AI Magazine (Raschka) Models Jun 17, 2025 👁 46

Understanding and Coding the KV Cache in LLMs from Scratch

KV caches are one of the most critical techniques for efficient inference in LLMs in production. KV caches…

Synced Review Research Jun 16, 2025 👁 41

MIT Researchers Unveil “SEAL”: A New Step Towards Self-Improving AI

The concept of AI self-improvement has been a hot topic in recent research circles, with a flurry of papers…

Hugging Face Open Source Jun 12, 2025 👁 49

How Long Prompts Block Other Requests - Optimizing LLM Performance

Back to Articles How Long Prompts Block Other Requests - Optimizing LLM Performance Team Article Published…

Air Street Capital Business Jun 8, 2025 👁 43

State of AI: June 2025 newsletter

Hi everyone!Welcome to the latest issue of the State of AI newsletter, an editorialized newsletter covering…

Hugging Face Open Source May 23, 2025 👁 40

Dell Enterprise Hub is all you need to build AI on premises

Back to Articles Dell Enterprise Hub is all you need to build AI on premises Published May 23, 2025 Update on…

Cameron Wolfe (AI) Models May 19, 2025 👁 44

A Guide for Debugging LLM Training Data

(from [2])Most discussions of LLM training focus heavily on models and algorithms. We enjoy experimenting…

Synced Review Research May 15, 2025 👁 43

DeepSeek-V3 New Paper is coming! Unveiling the Secrets of Low-Cost Large Model Training through Hardware-Aware Co-design

A newly released 14-page technical paper from the team behind DeepSeek-V3, with DeepSeek CEO Wenfeng Liang as…