Search AI News
Find articles from 100+ AI sources
Found 168 results for "PyTorch"
Low Precision Flash Attention 4: End-to-End Block-Scaled Attention for Blackwell
TL;DR We extend FlashAttention-4 [1] with MXFP8 forward and backward, reaching 2.85 PF/s forward and 2 PF/s…
MLPerf Inference v6.1: pioneering agent, VLM benchmarks
September 16, 2026 • 14 min read First agentic workload on datacenter hardware in MLPerf, and the first…
Announcing instance preference lists for Amazon SageMaker AI training jobs
Getting access to the right GPUs when you need them is one of the biggest challenges in training or…
[AINews] AEF-1 standard emerges for Third Party Evaluators, as Xai, OpenAI, and Anthropic all cosign
We last highlighted the pacing debate in July when Pacing the Frontier first emerged:And it seems that…
Helion x 🤗 HF Kernels: Building and Shipping Out-of-the-box Performant Kernels
TL;DR The HuggingFace Kernels project now has Helion support. This blog walks through how to build, autotune,…
Simplify and support your TorchServe workloads using Ray Serve Deep Learning Containers
TorchServe is no longer actively maintained. The official project notice states there are no planned updates,…
Pathway’s brain-inspired architecture development on Amazon SageMaker HyperPod
As AI systems take on more complex tasks, much of the industry’s progress has come from increasing model…
How SPREEAI trains the model behind photorealistic virtual try-on
September 8, 2026 • 23 min read A shopper uploads one photo. Ten seconds later, they’re looking at…
Build a Physical AI model factory with NVIDIA Cosmos 3 on SageMaker HyperPod
A Physical AI system, such as a robot or autonomous vehicle (AV) that translates real-world data into…
Run agent-driven Amazon SageMaker HyperPod operations with InstantStart
If you run foundation model (FM) workloads on Amazon SageMaker HyperPod, you know the work is rarely a single…
Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2
This post is a collaboration between AWS, NVIDIA and Heidi. Reducing automatic speech recognition (ASR)…
Bring your own model with Amazon SageMaker AI: Script mode in SDK v3
In 2021, we published Bring your own model with Amazon SageMaker script mode. That post showed how to use…
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
A few years ago, Caltech Prof. and co-founder of Accelerated Understanding, Anima Anandkumar set out to…
Granite 4.2 LLMs: How They're Built
Back to Articles Granite 4.2 LLMs: How They're Built Enterprise Article Published August 25, 2026 Upvote 30…
Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye
Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from…
Harnessing AI for Day-One Model Enablement
TL;DR The AI model landscape never stops moving, and the software stack that runs those models is always a…
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
Back to Articles Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers Published August…
[AINews] Cursor's $60B acquisition by SpaceXai closes
Throwback to when we did the first ever podcast on Cursor when they were 5 people:And then recapping agents…
Record, train, and deploy from one place with Strands Agents, LeRobot, and Hugging Face Storage Buckets
Back to Articles Record, train, and deploy from one place with Strands Agents, LeRobot, and Hugging Face…
FP8 Training on AMD GPUs with TorchTitan and TorchAO: Upstreaming Performance Improvements
At the PyTorch Conference 2025, we demonstrated linear scaling beyond 1,000 GPUs on AMD Instinct clusters…
[AINews] SpaceXAI Grok 4.6 and Grok @Bot
One of our top recurring themes of the year has been coding agents breaking containment into knowledge work,…
Best Self-Hosted Inference Servers for Open-Source Models: 7 Options Compared in 2026
From Ollama and LocalAI to vLLM and SIE, this guide matches each server to the workloads, model fleets, and…
Fast, On Device Agentic AI with Muse Glimmer on ExecuTorch
Today, Meta introduced Muse Glimmer, an open-weight, 30-billion-parameter model distilled from Meta’s Muse…
Making Knowledge Distillation Cheap Enough to Run at Scale
Back to Articles Making Knowledge Distillation Cheap Enough to Run at Scale Team Article Published August 10,…