Search AI News
Find articles from 100+ AI sources
Found 298 results for "LLaMA"
Deploying quantized models on Amazon SageMaker AI with Unsloth
This post was co-written with Daniel Han and Michael Han from Unsloth. Deploying large foundation models…
Disaggregated prefill and decode for LLM inference on SageMaker HyperPod
When prefill and decode share a GPU, long prompts stall token generation for every concurrent request.…
Towards Free Normalization: Fusing Normalization into GEMM and Attention Kernels
Code available at:…
[AINews] OpenAI launches GPT 5.6 Sol/Terra/Luna, Codex becomes ChatGPT superapp
On any other day, the launch of a surprisingly good/competitive Muse Spark 1.1 from Meta Superintelligence…
Ways to think about token pricing
There are only two things you can say with certainty about token prices: we’re in a supply crunch, and this…
[AINews] Lilian Weng summarizes 35 papers on Harness Engineering for RSI
Congrats to Meta Superintelligence on having the top 2/3 image/video models in the world! This would’ve…
Hugging Face Models on Foundry Managed Compute
Back to Articles Hugging Face Models on Foundry Managed Compute Enterprise Article Published July 7, 2026…
[AINews] The Field Guide to Fable
While we congratulate (friend of the show!) General Intuition on their new model and (friend of the show!)…
Bringing PyTorch Monarch to AMD GPUs: Single-Controller Distributed Training on ROCm
Training state-of-the-art large language models (LLMs) with billions of parameters requires distributed…
🔬 The Coolest Diffusion Research Isn't in LLMs — Evan Feinberg & Sergey Edunov, Genesis Molecular AI
This episode has a fun personal twist: There’s a counterfactual world where I was employee #1 at Genesis…
[AINews] Sonnet 5 today, and Fable 5 tomorrow
In separate announcements, Sonnet 5 was released today, and Fable/Mythos 5 were approved to be released again…
Ahmad Osman on why local AI is catching up
Ahmad Osman at the AI Engineer World’s Fair today.Ahmad Osman has been advocating for local AI — running…
Announcing transcribe.cpp
We are excited to announce that CJ Pais' transcribe.cpp has been officially released!What is…
Latest open artifacts (#22): Zyphra, Cohere, and Poolside are expanding the breadth of the ecosystem
A trend we continue to see in open model releases is that the ecosystem is becoming more diverse, with an…
Using Local Coding Agents
Many people reached out to me in the past asking about my local agent stack as well as how I set up my local…
Run a vLLM Server on HF Jobs in One Command
Back to Articles Run a vLLM Server on HF Jobs in One Command Published June 26, 2026 Update on GitHub Upvote…
GLM-5.2 is the step change for open agents
Housekeeping: Following my “State of the blog” post last week, noting a slight increase in paid features,…
The Hybrid AI Stack Is Coming for the Pricing Power of OpenAI and Anthropic
OpenAI and Anthropic are going public while still capturing much of the money spent on foundation-model…
MLPerf Training v6.0: Lambda delivers fastest LLM training on NVIDIA GB300 NVL72 and fastest MoE training on NVIDIA HGX B200
June 16, 2026 • 4 min read Lambda’s GB300 NVL72 Llama 3.1 8B MLPerf Training v6.0 submission improved…
Frontier post-training recipe review with Finbarr Timbers
As I’ve been recapping fundamentals of post-training to wrap up my RLHF / Post-training book I knew I…
Your AI bill is a tax on scale
Subscribe • Previous Issues The Hybrid AI Stack Is Coming for the Pricing Power of OpenAI and Anthropic…
DiffusionGemma: 4x faster text generation
Breadcrumb Innovation & AI Technology Developer tools DiffusionGemma: 4x faster text generation Jun 10, 2026…
NVIDIA Accelerates Google DeepMind’s DiffusionGemma for Local AI
Today, Google DeepMind released DiffusionGemma — an experimental open model built for exceptionally…
Introducing Gemma 4 12B: a unified, encoder-free multimodal model
Breadcrumb Innovation & AI Technology Developer tools Introducing Gemma 4 12B: a unified, encoder-free…