Search AI News
Find articles from 100+ AI sources
Found 207 results for "LLaMA"
How Long Prompts Block Other Requests - Optimizing LLM Performance
Back to Articles How Long Prompts Block Other Requests - Optimizing LLM Performance Team Article Published…
State of AI: June 2025 newsletter
Hi everyone!Welcome to the latest issue of the State of AI newsletter, an editorialized newsletter covering…
Dell Enterprise Hub is all you need to build AI on premises
Back to Articles Dell Enterprise Hub is all you need to build AI on premises Published May 23, 2025 Update on…
A Guide for Debugging LLM Training Data
(from [2])Most discussions of LLM training focus heavily on models and algorithms. We enjoy experimenting…
DeepSeek-V3 New Paper is coming! Unveiling the Secrets of Low-Cost Large Model Training through Hardware-Aware Co-design
A newly released 14-page technical paper from the team behind DeepSeek-V3, with DeepSeek CEO Wenfeng Liang as…
State of AI: May 2025 newsletter
Hi everyone!Welcome to the latest issue of the State of AI newsletter, an editorialized newsletter covering…
Coding LLMs from the Ground Up: A Complete Course
I wrote a lot about reasoning models in recent months (4 articles in a row)! Next to everything "agentic,"…
All About The Modern Positional Encodings In LLMs
The Positional Encoding in LLMs may appear somewhat mysterious the first time we come across the concept, and…
The State of Reinforcement Learning for LLM Reasoning
A lot has happened this month, especially with the releases of new flagship models like GPT-4.5 and Llama 4.…
Prefill and Decode for Concurrent Requests - Optimizing LLM Performance
Back to Articles Prefill and Decode for Concurrent Requests - Optimizing LLM Performance Team Article…
Introducing HELMET: Holistically Evaluating Long-context Language Models
Back to Articles Introducing HELMET: Holistically Evaluating Long-context Language Models Published April 16,…
Hugging Face to sell open-source robots thanks to Pollen Robotics acquisition 🤖
Back to Articles Hugging Face to sell open-source robots thanks to Pollen Robotics acquisition 🤖 Published…
4M Models Scanned: Protect AI + Hugging Face 6 Months In
Back to Articles 4M Models Scanned: Protect AI + Hugging Face 6 Months In Published April 14, 2025 Update on…
Defending against Prompt Injection with Structured Queries (StruQ) and Preference Optimization (SecAlign)
Recent advances in Large Language Models (LLMs) enable exciting LLM-integrated applications. However, as LLMs…
Arabic Leaderboards: Introducing Arabic Instruction Following, Updating AraGen, and More
Back to Articles Arabic Leaderboards: Introducing Arabic Instruction Following, Updating AraGen, and More…
LinkedIn Highlights, Mar 2025
Welcome to LinkedIn Highlights!Each month, I'll share my five seven top-performing LinkedIn posts, bringing…
State of AI: April 2025 newsletter
Hi everyone!Welcome to the latest issue of the State of AI newsletter, an editorialized newsletter covering…
The NLP Course is becoming the LLM Course
Back to Articles The NLP Course is becoming the LLM Course! Published April 3, 2025 Update on GitHub Upvote…
Vision Large Language Models (vLLMs)
After the popularization of text-based large language models (LLMs), one of the most important questions…
🚀 Accelerating LLM Inference with TGI on Intel Gaudi
Back to Articles 🚀 Accelerating LLM Inference with TGI on Intel Gaudi Published March 28, 2025 Update on…
Reduce AI Model Operational Costs With Quantization Techniques
Model quantization is becoming a core strategy for training and deployment! I am excited to introduce you to…
How To Improve Decoding Latency With Faster Self-Attention Mechanisms
In LLMs, handling large sequences is not enough, we need to make sure the decoding process is fast. Here we…
How To Reduce The Memory Usage Of The Self-Attention
With a bit of magic, we take a very inefficient computation like the Self-Attention and make it super…
LinkedIn Highlights, Feb 2025
Welcome to LinkedIn Highlights!Each month, I'll share my five six top-performing LinkedIn posts, bringing you…