Search AI News
Find articles from 100+ AI sources
Found 168 results for "PyTorch"
[AINews] Zawinski's Law of MultiAgents
We’ve discussed the HuggingFace-OpenAI security incident before, but OpenAI’s side of the story was the…
[AINews] AI is eating Finance; AIE NYC now open
We love writing a newsletter that cares more about being high signal than telling you there’s breaking news…
Specialized AI Is Getting Easier to Build
Last week I argued that open models will absorb most of the money and compute the world spends on AI. A week…
From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon
Figure 1: CUDA-to-MLX optimization translation map. CUDA optimization knowledge can be translated into…
The Big AI Labs Are Suddenly Competing with Your Own Data
Subscribe • Previous Issues Specialized AI Is Getting Easier to Build Last week I argued that open models…
[AINews] Much ado about Open Weights
Everyone say hi to Richard MacManus, our new Head of Editorial!The current debate about Open Weights is the…
Open Models and Open Weights Are Foundational to Secure AI
Home Blog Open Models and Open Weights Are Foundational to Secure AI 6 MIN READ Open Models and Open Weights…
Industry Leaders Unite in Open Secure AI Alliance for AI Safety and Security
Open source software is a critical pillar of the global economy. It underpins cloud computing, financial…
Build an explainable next-best-product recommendation system for banking on AWS
Building a deep learning-based explainable next-best-product recommendation system helps banking institutions…
Helion on TPU: Towards Hardware Heterogeneous Kernel Authoring
TL;DR Helion is PyTorch’s high-level DSL for writing performance-portable ML kernels. Partnering with…
OpenAI is scared of open-weight models. Should the US be?
The impressive capabilities of Chinese lab Moonshot’s Kimi K3, the biggest open-weight large language…
Fine-tune video and image models at scale with NVIDIA NeMo Automodel and 🤗 Diffusers
Back to Articles Fine-tune video and image models at scale with NVIDIA NeMo Automodel and 🤗 Diffusers…
Triton Plugin Extensions: Enabling TLX and Custom Compiler Passes Out of the Box
TLDR The PyTorch-Triton 3.7 release introduces the Triton Plugin Extensions system, a framework for…
Deploying quantized models on Amazon SageMaker AI with Unsloth
This post was co-written with Daniel Han and Michael Han from Unsloth. Deploying large foundation models…
Towards Free Normalization: Fusing Normalization into GEMM and Attention Kernels
Code available at:…
Import AI 464: Fables writes GPU kernels; AI automation; and analog computation
Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from…
🤗 Kernels: Major Updates
Back to Articles 🤗 Kernels: Major Updates Published July 6, 2026 Update on GitHub Upvote 26 +20 Sayak Paul…
Building the Future of On-Device AI at the ExecuTorch Hackathon
This past weekend in San Francisco, builders, researchers, mobile developers, and AI practitioners came…
TokenSpeed-Kernel: Portable APIs and High-Performance Kernels for Multi-Silicon LLM Inference
TL;DR The TokenSpeed-kernel is a standalone, open-source subsystem designed to solve backend complexity in…
Serving DeepSeek-V4 on GB300 with SGLang: 5x Higher Throughput at the Same Interactivity Since Day-0
TL;DR: DeepSeek-V4 support was live in SGLang on Day-0, but the Day-0 stack was only the starting point.…
MLPerf Training v6.0: Lambda delivers fastest LLM training on NVIDIA GB300 NVL72 and fastest MoE training on NVIDIA HGX B200
June 16, 2026 • 4 min read Lambda’s GB300 NVL72 Llama 3.1 8B MLPerf Training v6.0 submission improved…
The Open Source Community is backing OpenEnv for Agentic RL
Back to Articles The Open Source Community is backing OpenEnv for Agentic RL Published June 8, 2026 Update on…
Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention
After a short family break, I am excited to be back and catching up on a busy few weeks of open-weight LLM…
Unlocking asynchronicity in continuous batching
Back to Articles Unlocking asynchronicity in continuous batching Published May 14, 2026 Update on GitHub…