⚡ LIVE

Search AI News

Find articles from 100+ AI sources

🔍

Found 207 results for "LLaMA"

207 articles
Generative AI Pub Image AI 4w ago 👁 19

Your LLM Isn’t Thinking — It’s an Engineer Pulling Weights at 10,000 Tokens Per Second

What really happens inside an AI model, and why the engine running it matters more than you think.When most…

Generative AI Pub Image AI 4w ago 👁 21

The RAG Complexity Trap: Do More Components Actually Improve Retrieval Performance?

Modern RAG systems often include rerankers, hybrid retrieval, query rewriting, HNSW tuning, and many other…

Generative AI Pub Image AI 4w ago 👁 30

The Paragraph Buried in Anthropic’s July 1 Blog That Changes Every Enterprise AI Contract

Analysis | Hard InterruptUS Export control is all about control and imagined threatsOn July 1, 2026,…

TheSequence Business 4w ago 👁 18

The Sequence Radar #893: Last Week in AI: GPT-5.6, Grok 4.5, Muse Spark 1.1 and the Post-Chatbot Stack

Next Week in The Sequence:We continue our series about model distillation. In the AI of the Week, we are…

Latent Space Models 4w ago 👁 30

[AINews] not much happened today

So dancing bugs got upstaged by kpop girls, there’s the whole Bun vs Zig drama, and yesterday’s…

AWS ML Tools 4w ago 👁 14

Deploying quantized models on Amazon SageMaker AI with Unsloth

This post was co-written with Daniel Han and Michael Han from Unsloth. Deploying large foundation models…

AWS ML Tools 4w ago 👁 544

Disaggregated prefill and decode for LLM inference on SageMaker HyperPod

When prefill and decode share a GPU, long prompts stall token generation for every concurrent request.…

PyTorch Tools 4w ago 👁 15

Towards Free Normalization: Fusing Normalization into GEMM and Attention Kernels

Code available at:…

Latent Space Models 4w ago 👁 32

[AINews] OpenAI launches GPT 5.6 Sol/Terra/Luna, Codex becomes ChatGPT superapp

On any other day, the launch of a surprisingly good/competitive Muse Spark 1.1 from Meta Superintelligence…

Benedict Evans Business Jul 9, 2026 👁 23

Ways to think about token pricing

There are only two things you can say with certainty about token prices: we’re in a supply crunch, and this…

Latent Space Models Jul 8, 2026 👁 18

[AINews] Lilian Weng summarizes 35 papers on Harness Engineering for RSI

Congrats to Meta Superintelligence on having the top 2/3 image/video models in the world! This would’ve…

Hugging Face Open Source Jul 7, 2026 👁 18

Hugging Face Models on Foundry Managed Compute

Back to Articles Hugging Face Models on Foundry Managed Compute Enterprise Article Published July 7, 2026…

Latent Space Models Jul 7, 2026 👁 25

[AINews] The Field Guide to Fable

While we congratulate (friend of the show!) General Intuition on their new model and (friend of the show!)…

PyTorch Tools Jul 6, 2026 👁 16

Bringing PyTorch Monarch to AMD GPUs: Single-Controller Distributed Training on ROCm

Training state-of-the-art large language models (LLMs) with billions of parameters requires distributed…

Latent Space Models Jul 1, 2026 👁 22

🔬 The Coolest Diffusion Research Isn't in LLMs — Evan Feinberg & Sergey Edunov, Genesis Molecular AI

This episode has a fun personal twist: There’s a counterfactual world where I was employee #1 at Genesis…

Latent Space Models Jul 1, 2026 👁 20

[AINews] Sonnet 5 today, and Fable 5 tomorrow

In separate announcements, Sonnet 5 was released today, and Fable/Mythos 5 were approved to be released again…

Latent Space Models Jun 30, 2026 👁 1,076

Ahmad Osman on why local AI is catching up

Ahmad Osman at the AI Engineer World’s Fair today.Ahmad Osman has been advocating for local AI — running…

Mozilla AI Open Source Jun 30, 2026 👁 22

Announcing transcribe.cpp

We are excited to announce that CJ Pais' transcribe.cpp has been officially released!What is…

Interconnects AI Models Jun 28, 2026 👁 18

Latest open artifacts (#22): Zyphra, Cohere, and Poolside are expanding the breadth of the ecosystem

A trend we continue to see in open model releases is that the ecosystem is becoming more diverse, with an…

AI Magazine (Raschka) Models Jun 27, 2026 👁 23

Using Local Coding Agents

Many people reached out to me in the past asking about my local agent stack as well as how I set up my local…

Hugging Face Open Source Jun 26, 2026 👁 19

Run a vLLM Server on HF Jobs in One Command

Back to Articles Run a vLLM Server on HF Jobs in One Command Published June 26, 2026 Update on GitHub Upvote…

Interconnects AI Models Jun 22, 2026 👁 16

GLM-5.2 is the step change for open agents

Housekeeping: Following my “State of the blog” post last week, noting a slight increase in paid features,…

Gradient Flow Models Jun 17, 2026 👁 18

The Hybrid AI Stack Is Coming for the Pricing Power of OpenAI and Anthropic

OpenAI and Anthropic are going public while still capturing much of the money spent on foundation-model…

Lambda Labs Models Jun 16, 2026 👁 19

MLPerf Training v6.0: Lambda delivers fastest LLM training on NVIDIA GB300 NVL72 and fastest MoE training on NVIDIA HGX B200

June 16, 2026 • 4 min read Lambda’s GB300 NVL72 Llama 3.1 8B MLPerf Training v6.0 submission improved…