PROFILE
About
Focused on AI applications, Agent engineering, and production delivery.
I’m Liu ZhuoQi, an AI Application Engineer focused on Agent systems. I mainly build AI applications, Agent workflows, and production delivery systems.
I served as technical lead and led development for AISEO, Help Center, and UMS. I have also deployed and integrated SGLang / vLLM services for a video-understanding Agent. I am currently building Vane independently.
My proficiency varies across technologies. I care more about owning development, testing, troubleshooting, and deployment outcomes than presenting every framework I have used as an area of mastery.
Projects & Responsibilities#
AISEO
Built a multi-stage Agent content pipeline with Python / FastAPI / Temporal, covering RAG, multi-model routing, token costs, and multi-CMS publishing.Help Center
Developed a multi-site knowledge platform with Java / Quarkus / React, integrating RAG retrieval, vector storage, and local embeddings.UMS
Built a multi-product management platform with Go / Gin / Casdoor for shared authentication, permissions, subscription plans, and entitlement sync.Vane
Building with Go / PostgreSQL / Temporal and React / TypeScript, covering workflows, a Feishu Agent, and production deployment.Capabilities & Technologies#
AI Applications & Agents
Multi-stage Agent Workflows · RAG · Tool Use · Multi-Model Routing · Feedback Loops · Agent EvaluationModel Serving
SGLang · vLLM · Multimodal Video Understanding · GPU Integration and BenchmarkingBackend & Workflows
Python / FastAPI · Java / Quarkus · Go / Gin · Temporal · PostgreSQL · RedisFrontend
React · Vue 3 · TypeScript · ViteEngineering Delivery
Linux · Docker · Kubernetes · Jenkins · GitHub Actions · Ansible · Prometheus / GrafanaProfessional Certifications#
Technical Writing#
Why LLMs Can’t Remember You — Memory Mechanisms Dissected — 67 primary sources cross-validated across Anthropic / OpenAI / Google / Cursor docs and Karpathy / LeCun / Raschka papers, tearing down Agent memory systems from architectural constraints to product implementation.
How to Choose an LLM Inference Engine — A 2026 Map — 8 engines from vLLM / SGLang / TensorRT-LLM, plus PD disaggregation / speculative decoding / FP4 quantization, cross-checked against official blogs, GitHub, and arXiv.
Why We Migrated from Celery to Temporal — Workflow-engine selection for a production Agent pipeline, drawn from hitting each pitfall in the field rather than comparing docs.
An Agent Memory Selection Framework — Cost, latency, precision, and maintainability trade-offs across RAG / LLM Wiki / plain-text memory, with a decision tree.
Contact#
GitHub: github.com/YouToco
X / Twitter: x.com/busygod9527
Telegram: t.me/happyforyou0
WeChat: View QR Code
If you’re working on AI applications, Agent engineering, or the infrastructure around them, feel free to reach out through any channel above.