Skip to main content
TUR Tech Pulse — follow updates from the tools I use and recommend. Occasional, curated.

Content Tagged "evals"

This page collects all articles and projects tagged evals. Tags help you navigate by tool, theme, or topic—whether you're looking for content on specific technologies like n8n or Databricks, or broader themes like workflow automation or hockey analytics.

Content spans agentic engineering, production AI agents, data pipelines, full-stack work, and hockey analytics. Browse other tags, topics, or case studies for applied examples.

Projects describe case studies with real outcomes—conversational analytics prototypes, AI feedback agents, automated content pipelines, and hockey analytics dashboards. Blog posts cover implementation details, tool comparisons, and lessons learned. If you're looking for something specific, the blog index and portfolio offer alternative ways to explore.

Blog Posts

Automatic Model Watch: Know What New Models Are Good At Before You Route Traffic

field note ·
A practical pattern for watching new LLMs automatically — capability smoke tests, failure modes, and when to promote a model into production routing.

Eval Fixture Pack for Model Routing (Copy-Paste)

tutorial ·
A small, stable eval battery you can copy — fixtures for triage, JSON, edits, critique, and irreversible-action refusal — to promote models into agent lanes safely.

When Cheap Models Get Expensive: Pricing Failure Into Agent ROI

analysis ·
Token price is a trap. A free model that fails silently and burns human hours is expensive — here’s a simple way to price failure into agent ROI.