Resources
Proof and deep dives
How we moved the metric, and the thinking behind the systems we ship - case studies, whitepapers, the Insights blog, and the papers that built modern AI.
Rewiring how brands stay visible in the age of AI search
Search is being rebuilt around answers, not links. CogNerd optimizes brands for ChatGPT, Gemini, Claude and Perplexity - keeping organic reach from quietly disappearing into AI answers.
CogNerd · 3x Organic reach
Services are the new software, and the full-stack firm wins
The world spends six dollars on services for every one on software. AI is finally coming for that larger number. A field guide to the full-stack AI firm, the shift from selling tools to owning outcomes, and why Welzin was built for it.
Welzin · 6:1 Services vs software spend
A voice agent that answers, books, and closes the loop
A national home-services brand was losing calls to hold times and after-hours gaps. Welzin shipped a production voice agent that answers naturally, books appointments, and writes back to the CRM - with guardrails and full observability.
Confidential · 92% Calls resolved
Evaluating AI Systems in Production
A deep field guide to evaluating AI systems, centred on LLM-as-a-judge: how judges score, how they fail, how to design and validate one against human labels, and how to wire it into CI without fooling yourself.
July 19, 2026 · 40 min read
Retrieval-Augmented Generation in Production
The implementation guide for teams past the demo: chunking that respects document structure, hybrid retrieval and fusion, re-ranking economics, enforced citations, per-user permissions, latency and cost budgets, and a failure taxonomy that points at the stage to fix.
July 19, 2026 · 13 min read
From Pilot to Production
Most AI pilots do not fail on model quality; they stall on ownership, data access, integration, and an undefined bar for what would have counted as success. A field guide to designing pilots that can be promoted, and the 90-day sequence that gets one into production.
July 19, 2026 · 11 min read
Claude Fable 5: what it is, how good it is, and when to use it
Anthropic's most capable model is the first from its Mythos tier, a rung above Opus. The positioning, the specs that matter, the benchmarks, and the practitioner's rule for when its 2x price is actually worth paying.
2026-07-16 · 16 min read
Build a Karpathy-style LLM wiki in Obsidian: the implementation guide
Everyone publishes the concept. Almost nobody publishes the fields, the ingest contract, or the dispatcher. This is a working LLM wiki taken apart piece by piece - the permission model, the three frontmatter fields that carry the system, the five-step ingest, the invariant that makes automation safe, and the two things that actually break.
2026-07-16 · 23 min read
Evaluate Your Large Language Model
Large Language Models (LLMs) like GPT-4, LLaMA, and Claude now power critical real-world applications - from customer support bots that might accidentally...
2026-02-27 · 17 min read
Open Knowledge Format (OKF): An Open Specification for the LLM-Wiki Pattern
Standardised the LLM wiki - knowledge as plain files
Spec v0.1 · 2026 · 2026
PaperBanana: Automating Academic Illustration for AI Scientists
Agentic pipelines aimed at the research workflow itself
arXiv 2026 · 2026
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Open-sourced the reasoning-model recipe
Nature 2025 · 2025
Looking for something we have not written yet?
Tell us the problem. If we have comparable work we will point you at it - and if we have not, we will say so.
Talk to us
