Search AI News
Find articles from 100+ AI sources
Found 323 results for "NVIDIA"
State of AI: April 2025 newsletter
Hi everyone!Welcome to the latest issue of the State of AI newsletter, an editorialized newsletter covering…
Efficient Request Queueing – Optimizing LLM Performance
Back to Articles Efficient Request Queueing – Optimizing LLM Performance Team Article Published April 2,…
🚀 Accelerating LLM Inference with TGI on Intel Gaudi
Back to Articles 🚀 Accelerating LLM Inference with TGI on Intel Gaudi Published March 28, 2025 Update on…
Reduce AI Model Operational Costs With Quantization Techniques
Model quantization is becoming a core strategy for training and deployment! I am excited to introduce you to…
How To Construct Self-Attention Mechanisms For Arbitrary Long Sequences
With Gemini models having a 2M tokens context size and Claude having a 200K tokens context size while having…
State of AI: March 2025 newsletter
Hi everyone!Welcome to the latest issue of the State of AI newsletter, an editorialized newsletter covering…
State of AI: January 2025 newsletter
Dear readers,Welcome to the latest issue of the State of AI newsletter, an editorialized newsletter covering…
Introducing multi-backends (TRT-LLM, vLLM) support for Text Generation Inference
Back to Articles Introducing multi-backends (TRT-LLM, vLLM) support for Text Generation Inference Published…
Transformers.js v3: WebGPU Support, New Models & Tasks, and More…
Back to Articles Transformers.js v3: WebGPU Support, New Models & Tasks, and More… Published October 22,…
Accelerate 1.0.0
Back to Articles Accelerate 1.0.0 Published September 13, 2024 Update on GitHub Upvote 54 +48 Zachary Mueller…
Fine-tune FLUX.1 with an API
Replicate Blog Fine-tune FLUX.1 with an API Posted September 9, 2024 by zeke Info You can now fine-tune…
Fine-tune FLUX.1 to create images of yourself
Replicate Blog Fine-tune FLUX.1 to create images of yourself Posted August 30, 2024 by zeke Info Update (May…
Deploy Meta Llama 3.1 405B on Google Cloud Vertex AI
Back to Articles Deploy Meta Llama 3.1 405B on Google Cloud Vertex AI Published August 19, 2024 Update on…
Fine-tune FLUX.1 with your own images
Replicate Blog Fine-tune FLUX.1 with your own images Posted August 15, 2024 by deepfates Info You can now…
The Convergence of Proprietary and Open Source LLMs
A few months ago, I wrote about building LLM systems with self-hosted, open source models that beat…
Apple intelligence and AI maximalism
No-one outside Apple has really used any Apple Intelligence features yet. It won't launch until the autumn,…
Faster assisted generation support for Intel Gaudi
Back to Articles Faster assisted generation support for Intel Gaudi Published June 4, 2024 Update on GitHub…
Subscribe to Enterprise Hub with your AWS Account
Back to Articles Subscribe to Enterprise Hub with your AWS Account Published May 9, 2024 Update on GitHub…
Ways to think about AGI
The manuscript for ‘A Logic Named Joe’ In 1946, my grandfather, writing as ‘Murray Leinster’,…
GaLore: Advancing Large Model Training on Consumer-grade Hardware
Back to Articles GaLore: Advancing Large Model Training on Consumer-grade Hardware Published March 20, 2024…
Quanto: a PyTorch quantization backend for Optimum
Back to Articles Quanto: a PyTorch quantization backend for Optimum Published March 18, 2024 Update on GitHub…
Hugging Face and Google partner for open AI collaboration
Back to Articles Hugging Face and Google partner for open AI collaboration Published January 25, 2024 Update…
AMD + 🤗: Large Language Models Out-of-the-Box Acceleration with AMD GPU
Back to Articles AMD + 🤗: Large Language Models Out-of-the-Box Acceleration with AMD GPU Published…
How to use AnimateDiff in ComfyUI (vid2vid)
AnimateDiff is an open source technology released in July 2023 [Github][research paper]; it's one of the best…