Tested firsthand,
written to be understood.

AI Engineering · AI Coding · AI Productivity · Software Engineering

Recent articles

Subscribe via RSS ↗
AI Coding

Agent-Native Docs: 5 Documents That Turn Vibe Coding into an AI Team

Stop writing better prompts. Design the environment your AI works in. Five short documents act as the operating system for your agent team.

4 min read
AI Engineering

AI Agent Memory in Three Layers: Context, Session, and Persistent Memory

Even with a million tokens, agents forget yesterday. Memory needs its own design. Three layers keep you out of confusion.

5 min read
AI Engineering

AI Cyber Index: 3 Tests That Grade Vulnerability-Finding Agents

Security agents have a new exam. It grades a full loop: find the flaw, reproduce it, and patch it without breaking features.

4 min read
AI Engineering

Finance AI Models: Why Finance Needs Its Own Numbers and Citations Test

In finance, one wrong number is not just a typo. Here is why domain-tuned models exist and how finance RAG should be tested.

4 min read
AI Engineering

Gated Frontier Models: What Developers Should Use When Top Models Stay Closed

Three top models sit behind locks. Here is how to design with what is open.

4 min read
AI Engineering

Sonnet 5.5 Follow-Up: Why High Effort Beats Max

Sonnet max tops Terminal-Bench but burns the most tokens ever measured. Your default should be high.

4 min read
AI Engineering

Agents API Computer Use: What You Need Before You Hand a Browser to AI

The click matters less than the boundaries. Which sites are allowed, who signs in, and who verifies the result decide whether browser agents are safe.

6 min read
AI Engineering

AI Agent Security Guide: Why You Need Permissions Before Computer Use

AI used to only read. Now it runs things. Permission design comes before prompt design.

4 min read
AI Engineering

AI Model Routing in Practice: Start Cheap, Escalate When Stuck

Don't send everything to the priciest model. Classify, try, escalate, and log — four steps are enough.

4 min read
AI Engineering

AI Model Speed and Latency Guide: Smart but Slow Is Unusable

Tokens per second is only half the story. Time to first token and total task time decide how fast a model feels.

5 min read
AI Engineering

Claude Sonnet 5.5 vs Opus 5.5: Max Effort Is Often Overkill

Opus max 58, Sonnet max 56. A 2-point gap rarely justifies max effort every time.

5 min read
AI Engineering

Gemini 4 Argon Long-Horizon Agents: What 1M Output Tokens Change

The 1M number matters less than how long one work trajectory can stay alive. Here is what Gemini 4 Argon means for long-horizon agents.

5 min read
AI Engineering

GPT-6.1 Sol Multi-Agent Beta: Do More Subagents Really Work Better?

Three agents are not three times better. Multi-agent wins only when you can split work into independent pieces.

6 min read
AI Engineering

GPT-6.1 Sol vs GPT-6 Astra: Why Cost Per Task Beats the Best Model

GPT-6.1 Sol token prices are one-fifth of Astra. But the real question is not how cheap tokens are — it is what one finished task costs.

5 min read
AI Engineering

MCP vs Agents SDK vs WebMCP: Which Should You Learn First?

The three are not rivals. They play different roles. Split them into connection, runtime, and web exposure.

5 min read
AI Engineering

MiMo-V2.6-Pro vs GLM-5.3 vs Kimi K3: How to Pick an Open-Weight Model

The top three open-weight models sit at 46, 45, and 44 points. The scores look close, so start from whether you can actually use each model.

5 min read
AI Engineering

OpenAI Dots: From Chatbot to Always-On AI Agent

An AI that stops when the chat ends and an AI that keeps working between conversations are built differently. Here is the anatomy of the persistent agent behind Dots.

6 min read
AI Engineering

When RAG Fails, Don't Blame the Embeddings First: A 5-Stage Debug Guide

The culprit behind a RAG failure is often upstream of search. Start with parsing.

5 min read
AI Engineering

Terminal-Bench 4.0 Update: How to Read the New AI Scoreboard

The same model scores differently by harness and version. Here is the new scoreboard and how to read it correctly.

5 min read
AI Engineering

Meta Ray-Ban Display Web Apps: Web Developers Can Build AI Glasses Apps

Do you need a native SDK to build an AI glasses app? Meta now opens a path where HTML, CSS, JavaScript, and WebMCP are enough for glasses experiences.

8 min read
AI Coding

Android Studio BYOA: The IDE Becomes an Agent Runtime

You no longer pick an agent inside the IDE — you bring your own. Whoever owns the verification loop wins.

5 min read
AI Engineering

Claude Opus 5.5 Agent Loops: Why Loop Count Sets Your Bill

Input $4, output $20. But what matters more is how many calls your agent needs to finish one task.

5 min read
AI Engineering

NVIDIA AIPerf: Measure Serving with p95 and Concurrency, Not Averages

One tokens/sec number matters less than how p95 collapses as concurrent users grow. That is the real production metric for serving.

5 min read
AI Engineering

Qwen Code: Your Coding Agent Now Delegates to Other Agents

The setup is shifting from developer plus one AI to developer plus orchestrator plus specialist agents.

4 min read
AI Engineering

Agents API vs Your Own Agent Loop: Where to Delegate and What to Own

Should you write your own agent loop or use a managed harness? The real question is who owns context, retry, recovery, and state.

3 min read
AI Engineering

How to Read AI Benchmarks: Scores, Effort, Harness, and Cost Together

If Claude scores 58 and GPT scores 53, is Claude always better? Learn how to turn benchmark tables into real selection criteria.

4 min read
AI Engineering

AI Benchmark Reward Hacking: When Passing Tests Is Not Real Success

Passing tests and solving the problem are not the same. Reward hacking is becoming a core skill for anyone who assigns work to AI agents.

4 min read
AI Engineering

Benchmark Saturation Explained: Why AI Tests Expire When Scores Top Out

In a test where everyone scores 95, how much do 96 and 97 matter? Learn to spot the moment an AI benchmark saturates.

3 min read
AI Engineering

Why AI Benchmark Scores Change: Check the Test Version Before the Model

A model drops from 55 to 48 overnight. Did it get worse? Check the benchmark version first.

3 min read
AI Engineering

AI Cost per Successful Task: The Pricing Math That Beats Token Rates

Is a cheap model that fails three times really cheap? Divide API spend by successful tasks and your model choice changes.

4 min read
AI Productivity

Learning English with AI: Build a Speaking Routine That Sticks

AI helps most not by giving answers but by multiplying your reps: speak, get corrected, and try again.

3 min read
AI Coding

AI Game Development for Beginners: Build Games with Gemini and Flutter

Do you need months of syntax study before building a game? With AI, you can start with a tiny game and learn what you need along the way.

2 min read
AI Coding

Build Your First App with Public Data: A Vibe Coding Intro

Real data beats sample data. Your app feels alive the moment it shows something true — here's how to start with a public data API.

2 min read
AI Coding

Chrome Extensions with AI: Can You Build One in an Hour?

One browser habit you repeat daily can become an extension idea. Start with a tiny tool and build it with AI.

2 min read
AI Coding

Claude Code for Real Projects: From Chat AI to Coding Agent

AI you ask about code and AI that works inside your project need different habits. Here's Claude Code from an agent's perspective.

2 min read
AI Engineering

Claude Opus 5.5 1M Context: Is Its Memory Really Better?

Fitting 1M tokens and using 1M tokens well are different skills. Here's the long context of Claude Opus 5.5 from a project perspective.

4 min read
AI Engineering

Claude Opus 5.5 vs GPT-6 Astra: Real Differences Beyond Scores

Opus 5.5 scores 58 on the Intelligence Index, Astra 53. Simple on paper — real model selection is far messier.

4 min read
AI Productivity

CodeBridge Course Guide: From AI Basics to Coding, RAG, and Agents

Too many courses to pick from? Start from your goal. This guide connects CodeBridge courses by topic and level.

2 min read
AI Coding

Coding Agent Rankings: Why Pass@1, Cost per Task, and Time per Task Belong Together

What if an agent scores 2 points higher but costs twice as much and runs twice as slow? Here is how to read all three axes together.

3 min read
AI Coding

SWE-bench vs Terminal-Bench vs ProgramBench: What Coding Benchmarks Actually Measure

They are all called coding benchmarks, but the exam questions are completely different. You need to separate bug fixes, terminal work, and program rebuilds.

4 min read
AI Engineering

DeepSeek V4.1 Flash Agent Cost: Why Token Price Alone Misleads You

Does the cheapest model make the cheapest agent? DeepSeek V4.1 Flash shows why KV cache design changes how you should view real AI task costs.

4 min read
AI Productivity

Excel Automation with AI: Start Without Knowing How to Code

If you copy and clean the same Excel files every week, that is an automation candidate. Here is the most realistic way to start with AI.

2 min read
AI Engineering

Gemini 3.8 Live Voice AI: Why Running Tools During Conversation Matters

Natural voice AI needs more than a good voice. How you handle tool work while the user keeps talking decides the whole experience.

4 min read
AI Productivity

Is Gemini 4 Pro Released? How to Fact-Check AI Rumors with Official Sources

Even if you see posts claiming Gemini 4 Pro launched, no model ID means no product yet. Here is the fastest order for checking new AI model news.

4 min read
AI Productivity

Generative AI Content Basics: Where to Start with Text, Images, and Video

Good content starts with your idea and editing bar, not tool names. Here is the starter flow for making content with generative AI.

2 min read
Software Engineering

Git Version Control in Practice: Why It Matters More in the AI Coding Era

When AI edits many files at once, what changed and how to revert it matters more. A practical look at Git for real work.

2 min read
AI Engineering

GPT-5.6 Sol Model Routing: Why One Strong Model for Everything Is Wasteful

Using the strongest model for every request is not always best. The GPT-5.6 family shows how to split models by task difficulty.

4 min read
AI Engineering

GPT-6 Astra Computer Use: Why Finish the Job Is Harder Than Answering

Answering well and finishing real work are different skills. GPT-6 Astra shows why completion criteria matter in end-to-end AI tasks.

4 min read
AI Productivity

GPT-6 Sol vs Luna: Why You Do Not Always Need the Expensive Model

Sol is stronger and Luna is much cheaper. The point is not picking one — it is deciding where to switch models.

4 min read
AI Coding

Grok 4.7 Self-Verification in Coding: Can It Replace Tests?

A model saying it checked itself is not the same as tests passing. Grok 4.7 is a good moment to learn the limits of self-verification.

4 min read
AI Engineering

Hallucination vs Abstention: Is Saying I Don't Know a Weak Model?

Which is better: an AI that always answers confidently, or one that stops when unsure? Hallucination rate alone cannot tell you.

4 min read
AI Engineering

Internal Document AI Chatbot: How RAG Answers From Company Docs

How do you make AI read company rules and manuals? Take a light tour of the basic RAG structure behind an internal document chatbot.

4 min read
AI Engineering

Kimi K3 Open-Weight Frontier Model: Why 2.8T Parameters Are Not the Point

More interesting than model size is that you can work with the weights directly. Kimi K3 shows what an open-weight model really gives you.

4 min read
AI Engineering

Luna to Sol to Astra Model Routing: Stop Sending Everything to the Best Model

Can you protect quality without sending every request to the priciest model? Design escalation routing that promotes only on failure.

4 min read
AI Productivity

Meta Muse Image Editing: Why Editing Beats One-Shot Generation

Picking one great image matters less than safely steering it. Muse Image makes repeated, controlled edits the real skill.

4 min read
AI Engineering

Meta Muse Personal AI Agent: What Changes When AI Acts for You

Writing an email and sending it are totally different problems. Meta Muse shows why permissions and action limits define personal AI agents.

4 min read
AI Coding

Muse Spark 1.3 and Long Coding Tasks: Intermediate State Beats the First Prompt

The longer AI works, the less the first prompt matters. What keeps it on track is visible intermediate state. Here is what Muse Spark 1.3 teaches us about long coding tasks.

4 min read
AI Engineering

METR Time Horizon: How Many Hours of Work Can an AI Agent Do Alone?

Instead of a benchmark score, ask this: work that takes a human 30 minutes, 2 hours, or 8 hours — how far can an AI agent go alone? Here is METR Time Horizon, explained simply.

3 min read
AI Engineering

Mistral OCR 4.1 and RAG: Check Document Parsing Before You Blame Search

What if your RAG misses answers because of PDF parsing, not the search model? Use Mistral OCR 4.1 as a reason to inspect the first stage of document AI.

4 min read
AI Engineering

Robostral Navigate and Physical AI: What Changes When an LLM Moves a Robot?

You can re-ask a chatbot, but a robot's wrong move can mean a collision. See the core difference of Physical AI through Robostral Navigate.

4 min read
AI Coding

Model vs Harness in AI Coding: When the Environment Matters More

Same model, yet great on one project and broken on another. The difference may sit outside the model, in the working environment you gave it.

4 min read
AI Engineering

MoE Active Parameters: Is 6B Active Really a 6B Model?

A 125B model with only 6B activated — is it as light as a 6B model? Split the two numbers MoE readers confuse most.

5 min read
AI Engineering

1M Context Windows: Why Long Context Does Not Replace RAG

With a 1M-token context, should you just paste everything in? Compare long context and RAG as cost structures, not replacements.

3 min read
AI Engineering

Open-Weight Model Costs: Are GLM, Kimi, and Qwen Really Cheaper?

Open weights sound free, but running them is not. Compare hosted API cost and self-hosting cost inside one frame.

3 min read
AI Engineering

OpenAI Agents API: The Codex Harness Becomes an API

From model-call APIs to the agent runtime itself as an API. See how the Agents API changes the structure you build.

5 min read
AI Coding

ProgramBench: The Test Where AI Rebuilds a Whole Program From Scratch

With no source code at all, how far can an AI coding agent get? ProgramBench tests the ability to rebuild software.

3 min read
AI Engineering

Prompt Caching: When Does It Actually Save You AI API Money?

Is sending the same 100K tokens every time always the same price? Let us work out how the prompt cache changes the real cost structure.

3 min read
Software Engineering

Qt and QML: Building Cross-Platform GUI Apps With C++

If you have only written console programs in C++, Qt and QML are a great next step into apps with real screens. Here is a light tour of both.

2 min read
Software Engineering

REST APIs in Qt Apps: What Can You Build With Real Data?

Once a Qt app meets a REST API, it can handle real data like search, weather, and music. Here is the basic flow.

3 min read
AI Engineering

Qwen3.8 Omni Flash: Should You Merge Image, Voice, and Video Into One Model?

Would merging OCR, speech recognition, and video analysis into one model really make life easier? A look at multimodal design through Qwen3.8 Omni Flash.

4 min read
AI Engineering

Qwen4 Architecture: What Are QSA, Gated Residual, and PLE?

The Qwen4 shift is not just about adding parameters. Here is why attention, residual, and embedding are changing at the same time.

5 min read
AI Engineering

Classic RAG vs GraphRAG vs Agentic RAG: What Actually Differs?

RAG has many names now, but the starting point is the same. Start with how each one finds information and hands it to the model.

2 min read
AI Engineering

Reasoning High vs Max: Does Thinking Longer Always Answer Better?

What changes when you run a problem Max already solves on High? Match thinking strength to task difficulty, not model quality.

3 min read
AI Engineering

Recursive Self-Improvement: How Close Is AI to Improving Itself?

Can AI really build the next AI by itself? We separate what is already automated from what humans still own.

5 min read
AI Engineering

Sparse Attention Explained: The Cost of 1M-Token Context

If context gets 10x longer, does attention cost 10x more? We run the numbers on why sparse attention matters again.

4 min read
AI Coding

SWE-bench Multimodal v2: When AI Fixes UI Bugs from Screenshots

What if a GitHub issue is only a screenshot? We look at SWE-bench Multimodal v2, which adds vision to real coding tasks.

3 min read
AI Productivity

Using Multiple AI Tools: Should You Stick to Just One?

More subscriptions do not mean more output. Learn a simple way to split work across AI tools by task and compare them fairly.

3 min read
AI Coding

Vibe Coding for Web Development: From Idea to Deployment

Even when AI writes the code, websites still move through planning, building, version control, and deployment. Here is the full flow.

3 min read
AI Coding

What Is Vibe Coding? Starting with GitHub Copilot the Right Way

AI can generate code fast, but you still must understand and verify the result. Here is the mindset that connects vibe coding to real work.

2 min read
AI Engineering

What Is Graph Engineering? Mapping Complex AI Work as Nodes and Flows

When one straight flow is not enough, think in nodes and links. Here is a gentle intro to the graph view of AI work.

3 min read
AI Engineering

What Is Harness Engineering? What Comes After Using AI Well

Is picking a good model enough? We start with the harness, the working system around the model that decides real agent performance.

3 min read
AI Engineering

What Is Jev? AI That Decides Without Writing Sentences

Sometimes your app needs a decision, not a sentence. We unpack why Jev exists, what it does well, and where it falls short.

10 min read
AI Engineering

What Is Loop Engineering? How AI Agents Work by Repeating

Not one answer, but acting, checking, and retrying. We look at why loops are the basic structure of agents.

3 min read
AI Engineering

RAG vs Fine-Tuning: Which One Solves Your Problem?

Feeding outside material to a model and training it further solve different problems.

3 min read
AI Engineering

What Is RAG? From Retrieval to Answer in One Guide

How can a language model answer from documents it never memorized? We trace how RAG finds evidence and writes answers.

3 min read