Goblin
News
AI news by
promptgoblins.ai
|
News
About
News
About
Filtered by:
Inference
Clear
Titles
Summaries
6
Speculative Programmatic Tool Calling Technique Reduces LLM Agent Latency by Pre-Launching Tool Calls During Generation
Research
1
6d ago
6
Speculative Programmatic Tool Calling Technique Reduces LLM Agent Latency by Pre-Launching Tool Calls During Generation
Research
· 1 src · 6d ago
Discuss
8
NVIDIA Groq 3 LPX Enters Full Production with 3,400 Tokens/Second, First Deployed at Nebius
Updated Aug 25
Infra
6
6d ago
8
NVIDIA Groq 3 LPX Enters Full Production with 3,400 Tokens/Second, First Deployed at Nebius
Top
Upd Aug 25
Infra
· 6 srcs · 6d ago
Discuss
7
NVIDIA Vera Rubin NVL72 Claims 30x Throughput Per Megawatt Gain Over GB300 for Agentic AI
Infra
1
Aug 24
7
NVIDIA Vera Rubin NVL72 Claims 30x Throughput Per Megawatt Gain Over GB300 for Agentic AI
Infra
· 1 src · Aug 24
Discuss
7
Recirculation: Training-Free Inference Technique Cuts Gemma3 Perplexity 23% and Boosts Math Accuracy 21%
Research
1
Aug 21
7
Recirculation: Training-Free Inference Technique Cuts Gemma3 Perplexity 23% and Boosts Math Accuracy 21%
Research
· 1 src · Aug 21
Discuss
7
Cerebras Launches CS-4: Rack-Scale AI Server with WSE-3 Turbo Chips Claiming 30x GPU Token Throughput
Updated Aug 21
Products
6
Aug 21
7
Cerebras Launches CS-4: Rack-Scale AI Server with WSE-3 Turbo Chips Claiming 30x GPU Token Throughput
Upd Aug 21
Products
· 6 srcs · Aug 21
Discuss
6
NVIDIA Releases TensorRT Model Connect: Two-Command HuggingFace-to-TensorRT Inference in Public Preview
Products
1
Aug 19
6
NVIDIA Releases TensorRT Model Connect: Two-Command HuggingFace-to-TensorRT Inference in Public Preview
Products
· 1 src · Aug 19
Discuss
6
Test-Time Training Lets AI Models Adapt During Inference, Reshaping Memory and Compute Trade-offs
Research
1
Aug 18
6
Test-Time Training Lets AI Models Adapt During Inference, Reshaping Memory and Compute Trade-offs
Research
· 1 src · Aug 18
Discuss
8
DeepSeek Raises V4 API Prices Up to 1,100% with Peak/Off-Peak Structure, Effective August 16
Updated Aug 14
Markets
3
Aug 14
8
DeepSeek Raises V4 API Prices Up to 1,100% with Peak/Off-Peak Structure, Effective August 16
Upd Aug 14
Markets
· 3 srcs · Aug 14
Discuss
6
Sapiom Raises $35M Series A to Slash AI Token Costs 10x, Launching Model Router to Compete With OpenRouter
Updated Aug 6
Markets
2
Aug 6
6
Sapiom Raises $35M Series A to Slash AI Token Costs 10x, Launching Model Router to Compete With OpenRouter
Upd Aug 6
Markets
· 2 srcs · Aug 6
Discuss
6
Zero-Mem: Agent Memory Operations Without Any LLM Calls or Token Costs
Research
1
Aug 5
6
Zero-Mem: Agent Memory Operations Without Any LLM Calls or Token Costs
Research
· 1 src · Aug 5
Discuss
7
Liquid AI Launches LFM2.5-2.6B: On-Device Agentic Model That Rivals Models 4× Its Size
Updated Aug 5
Models
5
Aug 5
7
Liquid AI Launches LFM2.5-2.6B: On-Device Agentic Model That Rivals Models 4× Its Size
Top
Upd Aug 5
Models
· 5 srcs · Aug 5
Discuss
6
DeepSeek V4 Flash (304B) Runs in Production on Single AMD MI300X at 168 tok/s
Infra
1
Aug 4
6
DeepSeek V4 Flash (304B) Runs in Production on Single AMD MI300X at 168 tok/s
Infra
· 1 src · Aug 4
Discuss
6
DSpark Speculative Decoding Added to llama.cpp, Improving Draft Acceptance via Markov Head
Open Source
1
Jul 29
6
DSpark Speculative Decoding Added to llama.cpp, Improving Draft Acceptance via Markov Head
Open Source
· 1 src · Jul 29
Discuss
7
Kimi K3 (2.8T Params) Runs on 80x Consumer RTX 5090 GPUs — First Frontier Model with Zero HBM
Open Source
1
Jul 28
7
Kimi K3 (2.8T Params) Runs on 80x Consumer RTX 5090 GPUs — First Frontier Model with Zero HBM
Open Source
· 1 src · Jul 28
Discuss
6
Liquid AI Releases LFM2.5-Encoders: 3.7× Faster Than ModernBERT at Long Context on CPU
Models
1
Jul 28
6
Liquid AI Releases LFM2.5-Encoders: 3.7× Faster Than ModernBERT at Long Context on CPU
Models
· 1 src · Jul 28
Discuss
8
AMD's Helios Rack-Scale System Confirmed at Advancing AI 2026: MI455X Specs, Anthropic Partnership, and $1.4T Market Forecast
Updated Jul 23
Infra
3
Jul 23
8
AMD's Helios Rack-Scale System Confirmed at Advancing AI 2026: MI455X Specs, Anthropic Partnership, and $1.4T Market Forecast
Top
Upd Jul 23
Infra
· 3 srcs · Jul 23
Discuss
7
General Compute Secures $400M Inference-Chip-Backed Loan in First-of-Its-Kind Deal
Markets
1
Jul 17
7
General Compute Secures $400M Inference-Chip-Backed Loan in First-of-Its-Kind Deal
Markets
· 1 src · Jul 17
Discuss
7
Fireworks AI Raises $1.5B at $17.5B Valuation Amid Surge in Demand for Cheaper AI Models
Markets
1
Jul 17
7
Fireworks AI Raises $1.5B at $17.5B Valuation Amid Surge in Demand for Cheaper AI Models
Markets
· 1 src · Jul 17
Discuss
8
SambaNova Raises $1B at $11B Valuation, Lands JPMorgan Chase as On-Premises Inference Partner
Markets
1
Jul 8
8
SambaNova Raises $1B at $11B Valuation, Lands JPMorgan Chase as On-Premises Inference Partner
Top
Markets
· 1 src · Jul 8
Discuss
6
French Startup ZML Launches Free Multi-Chip AI Inference Server LLMD to Break Vendor Lock-In
Infra
1
Jul 8
6
French Startup ZML Launches Free Multi-Chip AI Inference Server LLMD to Break Vendor Lock-In
Infra
· 1 src · Jul 8
Discuss
7
UC Berkeley's Residual Context Diffusion Boosts Diffusion LLM Accuracy by 5–10 Points
Research
1
Jul 3
7
UC Berkeley's Residual Context Diffusion Boosts Diffusion LLM Accuracy by 5–10 Points
Research
· 1 src · Jul 3
Discuss
9
AI Chip Startup Etched Exits Stealth with $5B Valuation, $800M Raised, and $1B in Inference Cluster Orders
Markets
3
Jun 30
9
AI Chip Startup Etched Exits Stealth with $5B Valuation, $800M Raised, and $1B in Inference Cluster Orders
Top
Markets
· 3 srcs · Jun 30
Discuss
6
Wayfinder Router Offers Deterministic, Offline LLM Query Routing Without Model Calls
Open Source
1
Jun 28
6
Wayfinder Router Offers Deterministic, Offline LLM Query Routing Without Model Calls
Open Source
· 1 src · Jun 28
Discuss
7
Liquid AI Releases LFM2.5-230M: Tiny Edge Model Runs at 213 tok/s on Mobile and Powers Humanoid Robot On-Device
Models
1
Jun 26
7
Liquid AI Releases LFM2.5-230M: Tiny Edge Model Runs at 213 tok/s on Mobile and Powers Humanoid Robot On-Device
Models
· 1 src · Jun 26
Discuss
7
Unconventional AI Launches Oscillator-Based Architecture Targeting 1,000x Power Reduction for AI Inference
Infra
1
Jun 25
7
Unconventional AI Launches Oscillator-Based Architecture Targeting 1,000x Power Reduction for AI Inference
Infra
· 1 src · Jun 25
Discuss
6
Analysis: AI Model Parameters Could Reach 1.4 Quadrillion by 2031 Under Hardware and Data Constraints
Research
1
Jun 23
6
Analysis: AI Model Parameters Could Reach 1.4 Quadrillion by 2031 Under Hardware and Data Constraints
Research
· 1 src · Jun 23
Discuss
6
Sakana AI's AB-MCTS Algorithm Enables Multi-Model Collective Intelligence, Beats Individual Frontiers on ARC-AGI-2
Research
1
Jun 22
6
Sakana AI's AB-MCTS Algorithm Enables Multi-Model Collective Intelligence, Beats Individual Frontiers on ARC-AGI-2
Research
· 1 src · Jun 22
Discuss
6
Fireworks AI CEO Lin Qiao: Per-Application Custom Models Are the Key to Sustainable AI Agent Economics
Infra
1
Jun 18
6
Fireworks AI CEO Lin Qiao: Per-Application Custom Models Are the Key to Sustainable AI Agent Economics
Infra
· 1 src · Jun 18
Discuss
8
Google Releases DiffusionGemma: Open-Source 26B Model with 4x Faster Text Generation
Updated Aug 27
Models
45
5d ago
8
Google Releases DiffusionGemma: Open-Source 26B Model with 4x Faster Text Generation
Top
Upd Aug 27
Models
· 45 srcs · 5d ago
Discuss
6
FlashMemory: Lightweight Retriever Keeps DeepSeek-V4 KV-Cache at 10–15% GPU Footprint
Research
1
Jun 10
6
FlashMemory: Lightweight Retriever Keeps DeepSeek-V4 KV-Cache at 10–15% GPU Footprint
Research
· 1 src · Jun 10
Discuss
8
Xiaomi Breaks 1,000 Tokens/Sec Barrier on Trillion-Parameter Model Using Commodity GPUs
Updated Jun 15
Models
2
Jun 15
8
Xiaomi Breaks 1,000 Tokens/Sec Barrier on Trillion-Parameter Model Using Commodity GPUs
Top
Upd Jun 15
Models
· 2 srcs · Jun 15
Discuss
Filters
Signal
Title
Category
Sources
Posted
Discuss