Search AI News
Find articles from 100+ AI sources
Found 781 results for "LLM"
AI Agents from First Principles
(from [1] and source)The capabilities of large language models (LLMs) are advancing rapidly. As LLMs become…
State of AI: June 2025 newsletter
Hi everyone!Welcome to the latest issue of the State of AI newsletter, an editorialized newsletter covering…
LinkedIn Highlights, May 2025 - AI Coding Edition
Welcome to LinkedIn Highlights!Each month, I'll share my five seven top-performing LinkedIn posts, bringing…
AGI Is Not Multimodal
"In projecting language back as the model for thought, we lose sight of the tacit embodied understanding that…
KV Cache from scratch in nanoVLM
Back to Articles KV Cache from scratch in nanoVLM Published June 4, 2025 Update on GitHub Upvote 120 +114…
SmolVLA: Efficient Vision-Language-Action Model trained on Lerobot Community Data
HuggingFaceTB/SmolLM2-1.7B-Instruct Text Generation • 2B • Updated Apr 21, 2025 • 145k • 739
The Open-Source Toolkit for Building AI Agents v2
Welcome to a new post in the AI Agents Series - helping AI developers and researchers deploy and make sense…
GenAI’s adoption puzzle
This chart is very ‘glass half-empty or half-full?’, and it’s a puzzle. You could say that this is…
Dell Enterprise Hub is all you need to build AI on premises
Back to Articles Dell Enterprise Hub is all you need to build AI on premises Published May 23, 2025 Update on…
DeepSeek-V3 New Paper is coming! Unveiling the Secrets of Low-Cost Large Model Training through Hardware-Aware Co-design
A newly released 14-page technical paper from the team behind DeepSeek-V3, with DeepSeek CEO Wenfeng Liang as…
Falcon-Edge: A series of powerful, universal, fine-tunable 1.58bit language models.
BitNet Collection 🔥BitNet family of large language models (1-bit LLMs). • 7 items • Updated May 1,…
State of AI: May 2025 newsletter
Hi everyone!Welcome to the latest issue of the State of AI newsletter, an editorialized newsletter covering…
LinkedIn Highlights, Apr 2025
Welcome to LinkedIn Highlights!Each month, I'll share my five top-performing LinkedIn posts, bringing you the…
AGI is not a milestone
With the release of OpenAI’s latest model o3, there is renewed debate about whether Artificial General…
The 4 Things Qwen-3’s Chat Template Teaches Us
Back to Articles The 4 Things Qwen-3’s Chat Template Teaches Us Published April 30, 2025 Update on GitHub…
Sahar’s Coding with AI guide
Welcome to the first post in the AI Coding Series, where I'll share the strategies and insights I've…
PipelineRL
Back to Articles PipelineRL Enterprise Article Published April 25, 2025 Upvote 46 +40 Alex Piche alexpiche…
Join us for a Free LIVE Coding Event: Build The Self-Attention in PyTorch From Scratch
Next Friday, I am inviting you to join me for an exciting live coding event. It is a completely free event…
Can GRPO be 10x Efficient? Kwai AI’s SRPO Suggests Yes with SRPO
The remarkable success of OpenAI’s o1 series and DeepSeek-R1 has unequivocally demonstrated the power of…
17 Reasons Why Gradio Isn't Just Another UI Library
Back to Articles 17 Reasons Why Gradio Isn't Just Another UI Library Published April 16, 2025 Update on…
Introducing HELMET: Holistically Evaluating Long-context Language Models
Back to Articles Introducing HELMET: Holistically Evaluating Long-context Language Models Published April 16,…
DeepSeek Signals Next-Gen R2 Model, Unveils Novel Approach to Scaling Inference with SPCT
DeepSeek AI, a prominent player in the large language model arena, has recently published a research paper…
Visual Salamandra: Pushing the Boundaries of Multimodal Understanding
Back to Articles Visual Salamandra: Pushing the Boundaries of Multimodal Understanding Team Article Published…
Defending against Prompt Injection with Structured Queries (StruQ) and Preference Optimization (SecAlign)
Recent advances in Large Language Models (LLMs) enable exciting LLM-integrated applications. However, as LLMs…