Search AI News
Find articles from 100+ AI sources
Found 207 results for "LLaMA"
How to Beat Proprietary LLMs With Smaller Open Source Models
IntroductionWhen designing systems that use text generation models, many people first turn to proprietary…
A Guide to Structured Generation Using Constrained Decoding
IntroductionWe often want specific outputs when interacting with generative language models. This is…
Blazing Fast SetFit Inference with 🤗 Optimum Intel on Xeon
Back to Articles Blazing Fast SetFit Inference with 🤗 Optimum Intel on Xeon Published April 3, 2024 Update…
GaLore: Advancing Large Model Training on Consumer-grade Hardware
Back to Articles GaLore: Advancing Large Model Training on Consumer-grade Hardware Published March 20, 2024…
Quanto: a PyTorch quantization backend for Optimum
Back to Articles Quanto: a PyTorch quantization backend for Optimum Published March 18, 2024 Update on GitHub…
Car-GPT: Could LLMs finally make self-driving cars happen?
In 1928, London was in the middle of a terrible health crisis, devastated by bacterial diseases like…
Introducing the Red-Teaming Resistance Leaderboard
Back to Articles Introducing the Red-Teaming Resistance Leaderboard Published February 23, 2024 Update on…
AMD + 🤗: Large Language Models Out-of-the-Box Acceleration with AMD GPU
Back to Articles AMD + 🤗: Large Language Models Out-of-the-Box Acceleration with AMD GPU Published…
Optimum-NVIDIA Unlocking blazingly fast LLM inference in just 1 line of code
Back to Articles Optimum-NVIDIA on Hugging Face enables blazingly fast LLM inference in just 1 line of code…
The Artificiality of Alignment
This essay first appeared in Reboot. Credulous, breathless coverage of “AI existential risk” (abbreviated…
🧨 Accelerating Stable Diffusion XL Inference with JAX on Cloud TPU v5e
Back to Articles Accelerating Stable Diffusion XL Inference with JAX on Cloud TPU v5e Published October 3,…
Chat Templates: An End to the Silent Performance Killer
Back to Articles Chat Templates Published October 3, 2023 Update on GitHub Upvote 32 +26 Matthew Carrigan…
Run a Chatgpt-like Chatbot on a Single GPU with ROCm
Back to Articles Run a Chatgpt-like Chatbot on a Single GPU with ROCm Published May 15, 2023 Update on GitHub…
The Annotated Diffusion Model
Back to Articles The Annotated Diffusion Model Published June 7, 2022 Update on GitHub Upvote 365 +359 Niels…
How to generate text: using different decoding methods for language generation with Transformers
Back to Articles How to generate text: using different decoding methods for language generation with…