# ZhuoQi Dev > Liu ZhuoQi — AI application engineer focused on production Agent systems, backend engineering, context engineering, RAG, tool use, and AI security. This is Liu ZhuoQi's technical blog about AI agents, memory systems, context engineering, RAG, tool use, backend architecture and production practice. Every post exists in English and Chinese, and appending `index.md` to a post URL returns its Markdown source (e.g. /en/posts//index.md). Posts are listed newest first. ## Posts - [Agents That Check Their Own Work and Tidy Their Own Memory: Inside Claude Managed Agents Outcomes and Dreaming](https://zhuoqidev.com/en/posts/managed-agents-outcomes-dreaming/index.md): 2026-10-06 · Outcomes has an independent grader, blind to the writer's reasoning, check the work criterion by criterion and send it back until it passes; Dreaming reads an old memory store and past sessions between runs and writes a separate, reorganized store while leaving the original untouched. This article derives why both mechanisms are needed, shows the complete requests and events, walks through a real three-round revision, and translates both into contracts you can rebuild on any Agent stack. - [Why Codex Does Not Give the Model Every Tool: tool_search, BM25, and Model Replacement](https://zhuoqidev.com/en/posts/codex-first-request-tools/index.md): 2026-08-08 · A plain-language, source-grounded explanation of why Codex loads tools on demand, how BM25 ranks the best few, and how the same design transfers to Python, Go, and other models. - [Five Codex Harness Designs Worth Copying After Reading the Source](https://zhuoqidev.com/en/posts/codex-agent-design/index.md): 2026-08-02 · A source-grounded analysis of selective parallelism, Unified Exec, ExecCell, progressive Skills, persistent Goal continuation, and embedding the runtime through App Server. - [How Agents Remember You: Human Memory Science and a Code Audit of Six Open-Source Systems](https://zhuoqidev.com/en/posts/llm-memory-research/index.md): 2026-07-30 · From Ebbinghaus, H.M., working memory, and engrams to Mem0, Letta, Graphiti, LangMem, Cognee, and MemoryOS: a history of memory paradigms and a code-level comparison of what open-source agent memory systems actually implement. - [How to Choose an LLM Inference Engine — A 2026 Map from Local Single-GPU to PD Disaggregation](https://zhuoqidev.com/en/posts/llm-inference-engine-selection/index.md): 2026-07-19 · Aliyun's four-engine table — Ollama / vLLM / SGLang / HF Pipeline — is no longer enough in 2026. This piece re-maps inference engines into three tiers (Local → High-Performance Serving → Distributed/Disaggregated), covering 8 mainstream engines plus three new trends — PD disaggregation, speculative decoding, and FP4 quantization — with a decision matrix and a decision tree. - [OpenClaw in Practice: One File Path Eliminated 84% of Tool Calls — A Cron Job Debugging Story](https://zhuoqidev.com/en/posts/openclaw-cron-skill-optimization/index.md): 2026-06-20 · OpenClaw's daily-ai-news cron job kept timing out. The root cause wasn't a weak model, bloated prompt, or upstream API failure — it was a missing absolute path in the SKILL.md, causing the Agent to spend 15 exec calls searching for a tool's location every run. Message count ballooned from 54 to 165, tool results from 204KB to 1.1MB. This is the full debugging story. - [OpenClaw Memory in Practice: From 'Vector Search Is Down But Everything Still Works' to Zero-Cost NVIDIA Embeddings](https://zhuoqidev.com/en/posts/openclaw-memory-text-to-vector/index.md): 2026-06-20 · OpenClaw's vector retrieval silently failed — but BM25 text search kept the memory system running for two weeks unnoticed. When you discover 'it works without embeddings,' should you even bother fixing it? Here's how I used NVIDIA's free embedding API to complete the picture, and what I learned about when vector search actually matters. - [Claude's Tool Calling Paradigm Shift: A Deep Dive into Programmatic Tool Calling and Dynamic Filtering](https://zhuoqidev.com/en/posts/claude-programmatic-tool-calling-dynamic-filter/index.md): 2026-06-13 · Anthropic's Programmatic Tool Calling and Dynamic Filtering aren't just feature additions — they represent a paradigm shift in Agent architecture: from natural language orchestration to code-driven orchestration, from full context injection to on-demand filtering. Synthesizing multiple non-AI-written deep-dive articles, this post covers architecture, benchmarks, and production patterns. - [OpenClaw in Production: When the Most Advanced Memory System Meets the Quietest Failure](https://zhuoqidev.com/en/posts/openclaw-pitfalls/index.md): 2026-05-27 · A full-chain production battle log: from startup failures and Feishu message silent drops to production stability — compaction safeguard, five-layer debugging, model-harness fit, and memory system comparison. - [Why We Moved from Celery to Temporal for Production Agent Pipelines](https://zhuoqidev.com/en/posts/why-temporal-not-celery/index.md): 2026-05-16 · Agent pipelines are not ordinary async tasks — they have state, they get stuck, they need replay debugging, and one failure must not sink the entire batch. We hit every one of these walls with Celery before understanding that Temporal solves a fundamentally different problem. - [Where Do ChatGPT Business Promo Codes Come From? An OSINT Correction](https://zhuoqidev.com/en/posts/chatgpt-business-promo-investigation/index.md): 2026-05-14 · A source-by-source recheck of OpenAI, partner, Stripe, and linux.do evidence: public links, screened email codes, and partner-funded rebates still coexist. - [RAG vs LLM Wiki vs Plain Text — A Decision Framework for Agent Long-Term Memory](https://zhuoqidev.com/en/posts/memory-choice-framework/index.md): 2026-05-11 · Cost, latency, accuracy, and maintainability trade-offs across three Agent memory approaches: RAG, structured knowledge bases, and plain-text context memory. Includes decision tree. - [Embedding CSS Animation Demos in Hugo Articles](https://zhuoqidev.com/en/posts/css-animation-demo/index.md): 2026-05-04 · Use custom shortcodes to run live CSS animations directly in Hugo blog posts — no CodePen account needed. - [Building a Personal Site with Hugo and Dual-Stack CDN](https://zhuoqidev.com/en/posts/hugo-dual-cdn-blog/index.md): 2026-05-04 · How I set up Hugo + Blowfish with Alibaba Cloud OSS/CDN for China and Cloudflare Pages for international visitors — ICP filing, geo-DNS routing, and GitHub Actions dual-stack deployment. ## Projects - [Argus — A Browser-Local Long-Video Understanding Agent Harness](https://zhuoqidev.com/en/projects/argus/): An open-source, frontend-only agent harness for long-video understanding: video never leaves the browser, multi-provider LLMs via Vercel AI SDK, with frame-extraction / memory / sub-agent tools and dual-CDN publishing. Vite + React + TypeScript. - [Vane — A Production-Grade AI Personalized Feed System](https://zhuoqidev.com/en/projects/vane/): A self-built AI Agent system: multi-source ingestion → LLM personalized scoring → Feishu push, with a user-profile feedback loop. Go + Temporal + A2A protocol + React, full-stack. ## Optional - [All posts, full text](https://zhuoqidev.com/en/llms-full.txt) - [About the author](https://zhuoqidev.com/en/about/) - [中文版](https://zhuoqidev.com/llms.txt) - [RSS](https://zhuoqidev.com/en/index.xml)