Blog

中文

Start here

Build an agent, then choose its tools

Runnable code with real outputs: a 200-line agent loop with checkpoints, then the same task measured through MCP, CLI and Skills.

  1. 1.Python AI Agent Tutorial: Build an Agent That Searches, Calls Tools, and Keeps State
  2. 2.MCP, WebMCP, CLI, Skills: Which Tool Surface for the Same Task?

The geometry of learning

One thesis across four posts: deep learning is manifold geometry, and its sharpest failures are geometric tears you can measure.

  1. 1.The Four Realms of Neural Networks
  2. 2.DeepSeek V4 and Manifold Tearing
  3. 3.Chern's Question, Answered by Two Worlds at Once
  4. 4.How Torn Is a Trained Mixture-of-Experts?

Counting what you actually know

Three posts, one discipline: a memory that separates recorded, derived and guessed; a transfer test with shuffled controls; a factor library counted by independent bets, not by names.

  1. 1.AI Long-Term Memory: Similar Is Not the Same as Known
  2. 2.One Transfer Experiment: Bigger Fits Better, Transfers Worse
  3. 3.Factor Deduplication: Eleven Factors, Worth Fewer Than Five in Risk Terms (Python)

Engineering invariants

The incidents and invariants that separate production AI and quant systems from prototypes — one hard-won correctness rule per post.

  1. 1.Flatten Verifiers: When Your 'Flatten All' Order Doesn't Actually Flatten
  2. 2.Byte-Stability Tests for Prompt Caching
  3. 3.Mid-Turn Checkpointing in a Long-Running Agent Loop
  4. 4.Why Backtests Fail in Live Trading: A Cost and Fill Audit

September 26, 2026

AI Long-Term Memory: Similar Is Not the Same as Known

AI Research & Theory

Finding what a new thing resembles is easy; making it inherit what it should know is hard. In my structured-memory experiment, embedding scores recovered 2 of 24 facts a new entity should inherit; a written rule recovered all 24 and answered "no evidence" for all 40 facts the system had no record of. How to split an AI memory into lookup, derivation, and labelled guessing.

September 26, 2026

Factor Deduplication: Eleven Factors, Worth Fewer Than Five in Risk Terms (Python)

Dnalyaw Quant Trading

The number of factors in a library is not the number of independent signals. In one frozen sample, eleven momentum, value and size series built from Ken French public data have an average pairwise correlation of 0.13; grouping them into four correlation clusters leaves full-sample gross Sharpe essentially unchanged. The effective-number formula, a deduplication procedure, and a reproducible script.

September 26, 2026

One Transfer Experiment: Bigger Fits Better, Transfers Worse

AI Research & Theory

A high training score does not mean a model learned anything it can take elsewhere. In one transfer experiment on a memory model, quadrupling capacity raised the training score from 0.48 to 0.65 and lowered the score on unseen worlds from 0.26 to 0.21. Three checks that catch this: test on a new world, compare against a shuffled control, and audit for leaks.

September 17, 2026

MCP, WebMCP, CLI, Skills: Which Tool Surface for the Same Task?

AI Agent Practices

The same "close issues idle for 30 days" task handed to Claude through an MCP server, a bare command line, and the command line plus a SKILL.md, five runs each, measuring model calls, tool calls, token footprint, wall time and failure modes. All 15 runs were correct, but the cost differed by more than 2x: MCP averaged 3.2 model calls, CLI 7.4, Skills 6.0. With a reproducible single-file harness, and why WebMCP is not on the same track.

September 17, 2026

Historical Level 2 Data in Practice: Rebuilding an Order Book from MBO Messages in Python

Dnalyaw Quant Trading

You downloaded a sample of historical order book data. Now what? A dependency-free Python order book that replays 441,392 Nasdaq MBO messages for AAPL in 2.1 seconds, verified against the vendor MBP-10 feed at 148,402 events with zero mismatches, plus the MBO versus MBP distinction, the trade-record trap, and what you can measure once the book is rebuilt.

September 17, 2026

Python AI Agent Tutorial: Build an Agent That Searches, Calls Tools, and Keeps State

AI Agent Practices

A 200-line Python agent: web search, file tools, memory across runs, and resume-from-checkpoint after a crash. Install commands, the full code, output from four real runs with token counts, and one thing I did not expect: the search tool ships with a hidden sandbox.

August 19, 2026

DeepSeek Harness: a hot-swappable loop is not a production workflow

AI Agent Practices

DeepSeek Harness crossed six figures of GitHub stars in days and got framed as an open-source Claude Code killer. After the docs and the Cordis paper, it reads as a framework for recomposing agents. Kocoro is a local contract: Skill / MCP / Gateway already layered, the loop pinned. Those are not the same plugin kernel.

August 9, 2026

How Agents Run the Research Pipeline

Dnalyaw Quant Trading

Geometry of Alpha described the factory; One Person + AI described the org shape. This piece is the missing middle: how frozen gates become agent-executable contracts, what a candidate looks like walking the line overnight, which outputs are first-class (including every NO GO), and which doors must stay human — and fail closed — by construction.

August 9, 2026

Two Ways to Tame Discontinuity: K3 and DeepSeek V4

AI Research & Theory

Are Kimi K3 and DeepSeek V4 treating the same disease? A proposed unified reading—unbounded activations, discrete experts, cross-layer amplification—with different medicine. Stable LatentMoE, SiTU softcaps, and why MLA still sits in a hybrid stack.

July 6, 2026

One Person + AI: Rebuilding the Organizational Shape of a Quant Fund

Dnalyaw Quant Trading

A traditional quant fund scales research capacity with headcount: dozens of PhDs, each running a slice of the pipeline. Once the research process itself is frozen into an executable discipline, AI agents can take over enumeration, verification, monitoring, and reporting — and for the first time, throughput decouples from headcount. This piece covers the design principles of that new shape: what goes to the machine, what stays with the human, and why discipline is the precondition for AI, not the result of it.

July 5, 2026

Two Languages of Risk Control: Why Sizing and Safety Must Never Be Conflated

Dnalyaw Quant Trading

Position sizing answers "how big a bet." Safety control answers "when do you cut me off." These are two control systems with fundamentally different mathematical properties, and the most dangerous risk-design mistake in the industry is merging them into one. This piece covers where that line sits, why market-neutral does not mean crisis-neutral, and why real safety control needs veto power at the architecture level.

July 4, 2026

The Geometry of Alpha: Why the Quant Moat Is a Factory, Not a Recipe

Dnalyaw Quant Trading

The era of the single "magic signal" is over. The asset that actually compounds is not any one recipe — it is the factory that keeps producing, validating, and retiring signals. This piece is about the math behind weak-signal aggregation: why the √N law is the Pythagorean theorem, why factor stripping is a projection, why tail correlation is holonomy — and why reading markets and reading neural networks turn out to use the same geometry, the same four realms.

June 25, 2026

How Torn Is a Trained Mixture-of-Experts?

AI Research & Theory

I argued that loss spikes are geometric tears, not optimization blowups. So I went and measured the tear directly — in released MoE weights, OLMoE and Qwen. It is a genuine C⁰ jump at every layer, severe in the ~24% block-output jump, directional, and it survives training. A number behind the metaphor.

June 18, 2026

Chern's Question, Answered by Two Worlds at Once

AI Research & Theory

The manifold lineage of Riemann, Poincaré, and Chern passed down to Jiang Zehan, Boju Jiang, and Gen-hua Shi. Chern left a question behind — can piecewise-stacked manifolds be extended to arbitrarily complex regions? Two unrelated worlds, rock mechanics and neural networks, later arrived at the same answer.

May 29, 2026

DiffusionBlocks: The Network Already Knew It Was Three Parts

AI Research & Theory

Sakana AI's DiffusionBlocks trains a network in independent blocks and finds B=3 is the sweet spot. I argue why: the residual stream is already trifurcated into three geometric phases the manifold view predicts — so B=3 partitions a structure that functionally already exists.

May 27, 2026

Your AI Has a Brain. It Should Have a Heart.

AI Agent Practices

Models keep getting smarter while users keep doing more of the work — dragging in files, pasting emails, re-explaining who the client is. The missing piece is not a bigger context window. It is real memory. This is why I built Kocoro, a local, open-source Mac agent that remembers you.

May 5, 2026

DeepSeek V4 and Manifold Tearing

AI Research & Theory

Why training-time loss spikes are not optimization failures but geometric tears — DeepSeek V4 three defenses turn the "manifold prior" from an abstract stance into engineering evidence.

April 25, 2026

Transformer Architecture Book — English Edition Now Online

AI Research & Theory

The complete English edition of my Transformer book is now online — 32 chapters and 3 appendices that walk through every component of the Transformer, from tokenization through attention to a from-scratch implementation, and on to RLHF, Mixture-of-Experts, reasoning models, and post-Transformer architectures.

April 24, 2026

The Four Realms of Neural Networks

AI Research & Theory

A cultivation-novel reading of deep learning — PDE solvers, manifold geometry, gauge fields, and quantum attention — and why the last realm tells us AGI must be personal.

April 20, 2026

Why Backtests Fail in Live Trading: A Cost and Fill Audit

Dnalyaw Quant Trading

Separate signal failure from execution and accounting errors with a three-ledger comparison, a worked slippage example, and a reproducible commission-reconciliation test.

April 19, 2026

Mid-Turn Checkpointing in a Long-Running Agent Loop

AI Agent Practices

An agent turn can span 20 tool calls and 10 minutes. If the daemon dies at minute 9, naive designs throw everything away. A walk through why "just retry the turn" is wrong, the phase state machine that replaced it in Kocoro, and the checkpoint discipline that survives a SIGKILL mid-turn.

April 16, 2026

Byte-Stability Tests for Prompt Caching

AI Agent Practices

Prompt caching offers 90% cost savings when it works and 0% when it silently breaks. Why prompts are fragile in ways unit tests do not normally catch, and the testing discipline we built into Shannon to keep cache hit rates stable as the codebase evolves.

April 14, 2026

Flatten Verifiers: When Your 'Flatten All' Order Doesn't Actually Flatten

Dnalyaw Quant Trading

The single most safety-critical operation in any trading system is also one of the hardest correctness problems in execution engineering. A close-call incident, the race condition behind it, and the architectural pattern that replaced naive "close everything" logic with something correct by construction.

April 4, 2026

The AI Agent Harness: How Kocoro Evolved After Claude Code Went Public

AI Agent Practices

I built Kocoro — a Go agent runtime with tool dispatch, permissions, context management, and loop detection — before Claude Code went open source. When their architecture became public, the convergences were striking. Here is what we independently arrived at, what I learned from their caching practices, and what I think defines a production harness.

March 1, 2026

What Claude Code Learned About Multi-Agent Tool Design (and What Shannon Already Did)

AI Agent Practices

An Anthropic engineer described Claude Code's evolution from simple todo lists to dependency-aware task graphs — the same patterns Shannon independently arrived at. Five lessons mapped, including where Shannon needs to catch up.

February 17, 2026

Information Theory Is All You Need (to Understand LLMs)

AI Research & Theory

Shannon laid the foundation in 1948. Seventy-eight years later, his framework still explains why Transformers work, what cross-entropy loss actually means, and why your model can never be smarter than its training data.

February 14, 2026

Dnalyaw: Engineering an AI Quant Trading System from Scratch

Dnalyaw Quant Trading

Why vertical integration — research and execution inside a single unified pipeline — is the real moat in quant trading, and how Dnalyaw builds it with hundreds of features, disciplined risk, and a polyglot Rust/Go/Python execution core across global markets.

January 11, 2026

My Book on Building Production AI Agents

AI Agent Practices

A practical guide to building production-grade AI Agent systems. Covers single-agent design, multi-agent orchestration, MCP protocol, Computer Use, cost control, and enterprise deployment patterns.

January 11, 2026

My Book on Transformers and LLM Architecture

AI Research & Theory

A deep dive into every component of the Transformer architecture—from Tokenization to Attention mechanisms, from forward propagation to code implementation. For developers who want to truly understand how GPT and ChatGPT work.

January 6, 2026

2025 Year-In-Review and 2026 Prediction

Personal & Updates

Reflecting on how 2025 normalized agent workflows and reasoning models, and why 2026 feels less like a prediction and more like a state you can already opt into.

October 21, 2025

Tensor Logic: A Brain‑Like Architecture

AI Research & Theory

Bridging the gap between logic and learning—how tensor equations create AI systems that think both symbolically and intuitively.

October 7, 2025

Shannon: Designing a Production-Grade Multi-Agent Platform

AI Agent Practices

An architectural deep-dive into Shannon, a self-hosted multi-agent platform that addresses the three hardest problems in production AI: runaway costs, non-deterministic failures, and security vulnerabilities—through deliberate technology choices in Rust, Go, and Python.

April 29, 2025

AI Quantitative Trading: From Models to Quant Funds

Dnalyaw Quant Trading

Demystifying quantitative trading for AI practitioners — what quant funds actually do, why reinforcement learning fits markets better than LLMs, and where the real barriers lie.

February 25, 2025

From RNNs to LLMs: A Decade of Simplicity and Transformation

AI Research & Theory

Reflecting on Andrej Karpathy’s 2015 RNN post and the surprising evolution of LLMs with transformers and fine-tuning.

February 22, 2024

Transformer Architecture Explained: A Comprehensive Review

AI Research & Theory

A step-by-step walkthrough of the Transformer architecture: how the original encoder–decoder differs from decoder-only GPT, then input embeddings, positional encoding, self-attention, the FFN and output probabilities.