Search AI News
Find articles from 100+ AI sources
Found 781 results for "LLM"
Improving instruction hierarchy in frontier LLMs
IH-Challenge trains models to prioritize trusted instructions, improving instruction hierarchy, safety…
A Dream of Spring for Open-Weight LLMs: 10 Architectures from Jan-Feb 2026
If you have struggled a bit to keep up with open-weight model releases this month, this article should catch…
Alyah ⭐️: Toward Robust Evaluation of Emirati Dialect Capabilities in Arabic LLMs
Alyah ⭐️: Toward Robust Evaluation of Emirati Dialect Capabilities in Arabic LLMs. Hugging Face -…
Continual Learning with RL for LLMs
(from [1, 2, 3, 6, 11])Continual learning, which refers to the ability of an AI model to learn from new tasks…
Categories of Inference-Time Scaling for Improved LLM Reasoning
Inference scaling has become one of the most effective ways to improve answer quality and accuracy in…
The State Of LLMs 2025: Progress, Problems, and Predictions
As 2025 comes to a close, I want to look back at some of the year’s most important developments in large…
LLM Research Papers: The 2025 List (July to December)
In June, I shared a bonus article with my curated and bookmarked research paper lists to the paid subscribers…
AprielGuard: A Guardrail for Safety and Adversarial Robustness in Modern LLM Systems
AprielGuard: A Guardrail for Safety and Adversarial Robustness in Modern LLM Systems. Hugging Face -…
We Got Claude to Fine-Tune an Open Source LLM
We Got Claude to Fine-Tune an Open Source LLM. Hugging Face - artificial intelligence news.
Introducing AnyLanguageModel: One API for Local and Remote LLMs on Apple Platforms
Introducing AnyLanguageModel: One API for Local and Remote LLMs on Apple Platforms. Hugging Face - artificial…
Beyond Standard LLMs
From DeepSeek R1 to MiniMax-M2, the largest and most capable open-weight LLMs today remain autoregressive…
VaultGemma: The world's most capable differentially private LLM
Home Blog VaultGemma: The world's most capable differentially private LLM September 12, 2025Amer Sinha,…
Defining and evaluating political bias in LLMs
Learn how OpenAI evaluates political bias in ChatGPT through new real-world testing methods that improve…
Understanding the 4 Main Approaches to LLM Evaluation (From Scratch)
How do we actually evaluate LLMs?It’s a simple question, but one that tends to open up a much bigger…
AdalFlow: A PyTorch-Like Framework to Auto-Optimizing Prompt for your LLM agent
AI Agent frameworks are becoming just as important as model training itself! I am excited to introduce you to…
REINFORCE: Easy Online RL for LLMs
Reinforcement learning (RL) is playing an increasingly important role in research on large language models…
SyGra: The One-Stop Framework for Building Data for LLMs and SLMs
SyGra: The One-Stop Framework for Building Data for LLMs and SLMs. Hugging Face - artificial intelligence…
Fine-tune Any LLM from the Hugging Face Hub with Together AI
Fine-tune Any LLM from the Hugging Face Hub with Together AI. Hugging Face - artificial intelligence news.
Jupyter Agents: training LLMs to reason with notebooks
Jupyter Agents: training LLMs to reason with notebooks. Hugging Face - artificial intelligence news.
Online versus Offline RL for LLMs
(from [2, 5, 7, 9, 10])The alignment process teaches large language models (LLMs) how to generate completions…
Mixture-of-Experts: Early Sparse MoE Prototypes in LLMs
Mixture-of-Experts might be one of the most important improvements in the Transformer architecture! It allows…
Which Agent Causes Task Failures and When?Researchers from PSU and Duke explores automated failure attribution of LLM Multi-Agent Systems
Share My Research is Synced’s column that welcomes scholars to share their own research breakthroughs with…
🇵🇭 FilBench - Can LLMs Understand and Generate Filipino?
🇵🇭 FilBench - Can LLMs Understand and Generate Filipino?. Hugging Face - artificial intelligence news.
TextQuests: How Good are LLMs at Text-Based Video Games?
TextQuests: How Good are LLMs at Text-Based Video Games?. Hugging Face - artificial intelligence news.