The Local AI Architecture Upgrade: Ollama's Jev-Style Decision Models
Oct 02, 2026
00:00
04:42
ew episode 🎧 The Local AI Architecture Upgrade
Ollama 0.35 brings Jev-style decision models to your local machine. Instead of generating paragraphs, these small models make fast, strictly typed decisions and return clean JSON.
In this episode:
• What a Jev-style model is, and how the /v1/systemone endpoint works
• Nimble 9B (Bespoke Labs) and the Tev1 4B and 0.8B models (Together AI)
• ~91 ms per decision, fully local, with no cloud API costs
• Plugging it into C# .NET and WPF apps over a simple REST endpoint
• Using a tiny 0.8B model as a router in front of heavier models like Gemma 4 or Kimi
Try it: update to Ollama 0.35 and run `ollama pull nimble`.
How would you use structured decision models in your projects: data triage, or a router for your bigger LLMs? Let me know!
Narration voice and background music are AI-generated.
Thanks for the coffee ☕ It keeps these episodes coming.