BLOG
Tested firsthand,
written to be understood.
AI Engineering · AI Coding · AI Productivity · Software Engineering
Start here
FEATURED ARTICLEMeta Ray-Ban Display Web Apps: Web Developers Can Build AI Glasses Apps
Do you need a native SDK to build an AI glasses app? Meta now opens a path where HTML, CSS, JavaScript, and WebMCP are enough for glasses experiences.
Recent articles
Subscribe via RSS ↗Agent-Native Docs: 5 Documents That Turn Vibe Coding into an AI Team
Stop writing better prompts. Design the environment your AI works in. Five short documents act as the operating system for your agent team.
AI EngineeringAI Agent Memory in Three Layers: Context, Session, and Persistent Memory
Even with a million tokens, agents forget yesterday. Memory needs its own design. Three layers keep you out of confusion.
AI EngineeringAI Cyber Index: 3 Tests That Grade Vulnerability-Finding Agents
Security agents have a new exam. It grades a full loop: find the flaw, reproduce it, and patch it without breaking features.
AI EngineeringFinance AI Models: Why Finance Needs Its Own Numbers and Citations Test
In finance, one wrong number is not just a typo. Here is why domain-tuned models exist and how finance RAG should be tested.
AI EngineeringGated Frontier Models: What Developers Should Use When Top Models Stay Closed
Three top models sit behind locks. Here is how to design with what is open.
AI EngineeringSonnet 5.5 Follow-Up: Why High Effort Beats Max
Sonnet max tops Terminal-Bench but burns the most tokens ever measured. Your default should be high.
AI EngineeringAgents API Computer Use: What You Need Before You Hand a Browser to AI
The click matters less than the boundaries. Which sites are allowed, who signs in, and who verifies the result decide whether browser agents are safe.
AI EngineeringAI Agent Security Guide: Why You Need Permissions Before Computer Use
AI used to only read. Now it runs things. Permission design comes before prompt design.
AI EngineeringAI Model Routing in Practice: Start Cheap, Escalate When Stuck
Don't send everything to the priciest model. Classify, try, escalate, and log — four steps are enough.
AI EngineeringAI Model Speed and Latency Guide: Smart but Slow Is Unusable
Tokens per second is only half the story. Time to first token and total task time decide how fast a model feels.
AI EngineeringClaude Sonnet 5.5 vs Opus 5.5: Max Effort Is Often Overkill
Opus max 58, Sonnet max 56. A 2-point gap rarely justifies max effort every time.
AI EngineeringGemini 4 Argon Long-Horizon Agents: What 1M Output Tokens Change
The 1M number matters less than how long one work trajectory can stay alive. Here is what Gemini 4 Argon means for long-horizon agents.
AI EngineeringGPT-6.1 Sol Multi-Agent Beta: Do More Subagents Really Work Better?
Three agents are not three times better. Multi-agent wins only when you can split work into independent pieces.
AI EngineeringGPT-6.1 Sol vs GPT-6 Astra: Why Cost Per Task Beats the Best Model
GPT-6.1 Sol token prices are one-fifth of Astra. But the real question is not how cheap tokens are — it is what one finished task costs.
AI EngineeringMCP vs Agents SDK vs WebMCP: Which Should You Learn First?
The three are not rivals. They play different roles. Split them into connection, runtime, and web exposure.
AI EngineeringMiMo-V2.6-Pro vs GLM-5.3 vs Kimi K3: How to Pick an Open-Weight Model
The top three open-weight models sit at 46, 45, and 44 points. The scores look close, so start from whether you can actually use each model.
AI EngineeringOpenAI Dots: From Chatbot to Always-On AI Agent
An AI that stops when the chat ends and an AI that keeps working between conversations are built differently. Here is the anatomy of the persistent agent behind Dots.
AI EngineeringWhen RAG Fails, Don't Blame the Embeddings First: A 5-Stage Debug Guide
The culprit behind a RAG failure is often upstream of search. Start with parsing.
AI EngineeringTerminal-Bench 4.0 Update: How to Read the New AI Scoreboard
The same model scores differently by harness and version. Here is the new scoreboard and how to read it correctly.
AI EngineeringMeta Ray-Ban Display Web Apps: Web Developers Can Build AI Glasses Apps
Do you need a native SDK to build an AI glasses app? Meta now opens a path where HTML, CSS, JavaScript, and WebMCP are enough for glasses experiences.
AI CodingAndroid Studio BYOA: The IDE Becomes an Agent Runtime
You no longer pick an agent inside the IDE — you bring your own. Whoever owns the verification loop wins.
AI EngineeringClaude Opus 5.5 Agent Loops: Why Loop Count Sets Your Bill
Input $4, output $20. But what matters more is how many calls your agent needs to finish one task.
AI EngineeringNVIDIA AIPerf: Measure Serving with p95 and Concurrency, Not Averages
One tokens/sec number matters less than how p95 collapses as concurrent users grow. That is the real production metric for serving.
AI EngineeringQwen Code: Your Coding Agent Now Delegates to Other Agents
The setup is shifting from developer plus one AI to developer plus orchestrator plus specialist agents.
AI EngineeringAgents API vs Your Own Agent Loop: Where to Delegate and What to Own
Should you write your own agent loop or use a managed harness? The real question is who owns context, retry, recovery, and state.
AI EngineeringHow to Read AI Benchmarks: Scores, Effort, Harness, and Cost Together
If Claude scores 58 and GPT scores 53, is Claude always better? Learn how to turn benchmark tables into real selection criteria.
AI EngineeringAI Benchmark Reward Hacking: When Passing Tests Is Not Real Success
Passing tests and solving the problem are not the same. Reward hacking is becoming a core skill for anyone who assigns work to AI agents.
AI EngineeringBenchmark Saturation Explained: Why AI Tests Expire When Scores Top Out
In a test where everyone scores 95, how much do 96 and 97 matter? Learn to spot the moment an AI benchmark saturates.
AI EngineeringWhy AI Benchmark Scores Change: Check the Test Version Before the Model
A model drops from 55 to 48 overnight. Did it get worse? Check the benchmark version first.
AI EngineeringAI Cost per Successful Task: The Pricing Math That Beats Token Rates
Is a cheap model that fails three times really cheap? Divide API spend by successful tasks and your model choice changes.
AI ProductivityLearning English with AI: Build a Speaking Routine That Sticks
AI helps most not by giving answers but by multiplying your reps: speak, get corrected, and try again.
AI CodingAI Game Development for Beginners: Build Games with Gemini and Flutter
Do you need months of syntax study before building a game? With AI, you can start with a tiny game and learn what you need along the way.
AI CodingBuild Your First App with Public Data: A Vibe Coding Intro
Real data beats sample data. Your app feels alive the moment it shows something true — here's how to start with a public data API.
AI CodingChrome Extensions with AI: Can You Build One in an Hour?
One browser habit you repeat daily can become an extension idea. Start with a tiny tool and build it with AI.
AI CodingClaude Code for Real Projects: From Chat AI to Coding Agent
AI you ask about code and AI that works inside your project need different habits. Here's Claude Code from an agent's perspective.
AI EngineeringClaude Opus 5.5 1M Context: Is Its Memory Really Better?
Fitting 1M tokens and using 1M tokens well are different skills. Here's the long context of Claude Opus 5.5 from a project perspective.
AI EngineeringClaude Opus 5.5 vs GPT-6 Astra: Real Differences Beyond Scores
Opus 5.5 scores 58 on the Intelligence Index, Astra 53. Simple on paper — real model selection is far messier.
AI ProductivityCodeBridge Course Guide: From AI Basics to Coding, RAG, and Agents
Too many courses to pick from? Start from your goal. This guide connects CodeBridge courses by topic and level.
AI CodingCoding Agent Rankings: Why Pass@1, Cost per Task, and Time per Task Belong Together
What if an agent scores 2 points higher but costs twice as much and runs twice as slow? Here is how to read all three axes together.
AI CodingSWE-bench vs Terminal-Bench vs ProgramBench: What Coding Benchmarks Actually Measure
They are all called coding benchmarks, but the exam questions are completely different. You need to separate bug fixes, terminal work, and program rebuilds.
AI EngineeringDeepSeek V4.1 Flash Agent Cost: Why Token Price Alone Misleads You
Does the cheapest model make the cheapest agent? DeepSeek V4.1 Flash shows why KV cache design changes how you should view real AI task costs.
AI ProductivityExcel Automation with AI: Start Without Knowing How to Code
If you copy and clean the same Excel files every week, that is an automation candidate. Here is the most realistic way to start with AI.
AI EngineeringGemini 3.8 Live Voice AI: Why Running Tools During Conversation Matters
Natural voice AI needs more than a good voice. How you handle tool work while the user keeps talking decides the whole experience.
AI ProductivityIs Gemini 4 Pro Released? How to Fact-Check AI Rumors with Official Sources
Even if you see posts claiming Gemini 4 Pro launched, no model ID means no product yet. Here is the fastest order for checking new AI model news.
AI ProductivityGenerative AI Content Basics: Where to Start with Text, Images, and Video
Good content starts with your idea and editing bar, not tool names. Here is the starter flow for making content with generative AI.
Software EngineeringGit Version Control in Practice: Why It Matters More in the AI Coding Era
When AI edits many files at once, what changed and how to revert it matters more. A practical look at Git for real work.
AI EngineeringGPT-5.6 Sol Model Routing: Why One Strong Model for Everything Is Wasteful
Using the strongest model for every request is not always best. The GPT-5.6 family shows how to split models by task difficulty.
AI EngineeringGPT-6 Astra Computer Use: Why Finish the Job Is Harder Than Answering
Answering well and finishing real work are different skills. GPT-6 Astra shows why completion criteria matter in end-to-end AI tasks.
AI ProductivityGPT-6 Sol vs Luna: Why You Do Not Always Need the Expensive Model
Sol is stronger and Luna is much cheaper. The point is not picking one — it is deciding where to switch models.
AI CodingGrok 4.7 Self-Verification in Coding: Can It Replace Tests?
A model saying it checked itself is not the same as tests passing. Grok 4.7 is a good moment to learn the limits of self-verification.
AI EngineeringHallucination vs Abstention: Is Saying I Don't Know a Weak Model?
Which is better: an AI that always answers confidently, or one that stops when unsure? Hallucination rate alone cannot tell you.
AI EngineeringInternal Document AI Chatbot: How RAG Answers From Company Docs
How do you make AI read company rules and manuals? Take a light tour of the basic RAG structure behind an internal document chatbot.
AI EngineeringKimi K3 Open-Weight Frontier Model: Why 2.8T Parameters Are Not the Point
More interesting than model size is that you can work with the weights directly. Kimi K3 shows what an open-weight model really gives you.
AI EngineeringLuna to Sol to Astra Model Routing: Stop Sending Everything to the Best Model
Can you protect quality without sending every request to the priciest model? Design escalation routing that promotes only on failure.
AI ProductivityMeta Muse Image Editing: Why Editing Beats One-Shot Generation
Picking one great image matters less than safely steering it. Muse Image makes repeated, controlled edits the real skill.
AI EngineeringMeta Muse Personal AI Agent: What Changes When AI Acts for You
Writing an email and sending it are totally different problems. Meta Muse shows why permissions and action limits define personal AI agents.
AI CodingMuse Spark 1.3 and Long Coding Tasks: Intermediate State Beats the First Prompt
The longer AI works, the less the first prompt matters. What keeps it on track is visible intermediate state. Here is what Muse Spark 1.3 teaches us about long coding tasks.
AI EngineeringMETR Time Horizon: How Many Hours of Work Can an AI Agent Do Alone?
Instead of a benchmark score, ask this: work that takes a human 30 minutes, 2 hours, or 8 hours — how far can an AI agent go alone? Here is METR Time Horizon, explained simply.
AI EngineeringMistral OCR 4.1 and RAG: Check Document Parsing Before You Blame Search
What if your RAG misses answers because of PDF parsing, not the search model? Use Mistral OCR 4.1 as a reason to inspect the first stage of document AI.
AI EngineeringRobostral Navigate and Physical AI: What Changes When an LLM Moves a Robot?
You can re-ask a chatbot, but a robot's wrong move can mean a collision. See the core difference of Physical AI through Robostral Navigate.
AI CodingModel vs Harness in AI Coding: When the Environment Matters More
Same model, yet great on one project and broken on another. The difference may sit outside the model, in the working environment you gave it.
AI EngineeringMoE Active Parameters: Is 6B Active Really a 6B Model?
A 125B model with only 6B activated — is it as light as a 6B model? Split the two numbers MoE readers confuse most.
AI Engineering1M Context Windows: Why Long Context Does Not Replace RAG
With a 1M-token context, should you just paste everything in? Compare long context and RAG as cost structures, not replacements.
AI EngineeringOpen-Weight Model Costs: Are GLM, Kimi, and Qwen Really Cheaper?
Open weights sound free, but running them is not. Compare hosted API cost and self-hosting cost inside one frame.
AI EngineeringOpenAI Agents API: The Codex Harness Becomes an API
From model-call APIs to the agent runtime itself as an API. See how the Agents API changes the structure you build.
AI CodingProgramBench: The Test Where AI Rebuilds a Whole Program From Scratch
With no source code at all, how far can an AI coding agent get? ProgramBench tests the ability to rebuild software.
AI EngineeringPrompt Caching: When Does It Actually Save You AI API Money?
Is sending the same 100K tokens every time always the same price? Let us work out how the prompt cache changes the real cost structure.
Software EngineeringQt and QML: Building Cross-Platform GUI Apps With C++
If you have only written console programs in C++, Qt and QML are a great next step into apps with real screens. Here is a light tour of both.
Software EngineeringREST APIs in Qt Apps: What Can You Build With Real Data?
Once a Qt app meets a REST API, it can handle real data like search, weather, and music. Here is the basic flow.
AI EngineeringQwen3.8 Omni Flash: Should You Merge Image, Voice, and Video Into One Model?
Would merging OCR, speech recognition, and video analysis into one model really make life easier? A look at multimodal design through Qwen3.8 Omni Flash.
AI EngineeringQwen4 Architecture: What Are QSA, Gated Residual, and PLE?
The Qwen4 shift is not just about adding parameters. Here is why attention, residual, and embedding are changing at the same time.
AI EngineeringClassic RAG vs GraphRAG vs Agentic RAG: What Actually Differs?
RAG has many names now, but the starting point is the same. Start with how each one finds information and hands it to the model.
AI EngineeringReasoning High vs Max: Does Thinking Longer Always Answer Better?
What changes when you run a problem Max already solves on High? Match thinking strength to task difficulty, not model quality.
AI EngineeringRecursive Self-Improvement: How Close Is AI to Improving Itself?
Can AI really build the next AI by itself? We separate what is already automated from what humans still own.
AI EngineeringSparse Attention Explained: The Cost of 1M-Token Context
If context gets 10x longer, does attention cost 10x more? We run the numbers on why sparse attention matters again.
AI CodingSWE-bench Multimodal v2: When AI Fixes UI Bugs from Screenshots
What if a GitHub issue is only a screenshot? We look at SWE-bench Multimodal v2, which adds vision to real coding tasks.
AI ProductivityUsing Multiple AI Tools: Should You Stick to Just One?
More subscriptions do not mean more output. Learn a simple way to split work across AI tools by task and compare them fairly.
AI CodingVibe Coding for Web Development: From Idea to Deployment
Even when AI writes the code, websites still move through planning, building, version control, and deployment. Here is the full flow.
AI CodingWhat Is Vibe Coding? Starting with GitHub Copilot the Right Way
AI can generate code fast, but you still must understand and verify the result. Here is the mindset that connects vibe coding to real work.
AI EngineeringWhat Is Graph Engineering? Mapping Complex AI Work as Nodes and Flows
When one straight flow is not enough, think in nodes and links. Here is a gentle intro to the graph view of AI work.
AI EngineeringWhat Is Harness Engineering? What Comes After Using AI Well
Is picking a good model enough? We start with the harness, the working system around the model that decides real agent performance.
AI EngineeringWhat Is Jev? AI That Decides Without Writing Sentences
Sometimes your app needs a decision, not a sentence. We unpack why Jev exists, what it does well, and where it falls short.
AI EngineeringWhat Is Loop Engineering? How AI Agents Work by Repeating
Not one answer, but acting, checking, and retrying. We look at why loops are the basic structure of agents.
AI EngineeringRAG vs Fine-Tuning: Which One Solves Your Problem?
Feeding outside material to a model and training it further solve different problems.
AI EngineeringWhat Is RAG? From Retrieval to Answer in One Guide
How can a language model answer from documents it never memorized? We trace how RAG finds evidence and writes answers.
Nothing published in this category yet.