LLM engineer working on inference performance, agent safety, and agentic tooling.
I benchmark LLM serving on real (free-tier) GPUs and report measured numbers, fine-tune small guard models for agent safety, and build agents and operators that run in production-like setups. I also work on ServiceNow/ITSM automation and incident intelligence.
| Project | What it shows |
|---|---|
| inference-engineering | vLLM labs on a free T4 with measured results: 7.5× prefill speedup from prefix caching, 31× shorter ITL freeze with chunked prefill, throughput knee/cliff analysis, AWQ vs fp16, KV-cache capacity and preemption. |
| RocketGuard-1B | MiniCPM5-1B fine-tuned (Unsloth, ~48.7k examples) to make guardrail decisions for agents: allow / block / rewrite / ask. Weights on Hugging Face. |
| gemini-claw | Telegram-native Gemini CLI operator: private allowlisted chats, background task workers, session continuity, stream-JSON parsing, and safety controls. Tested with Vitest + CI. |
| ADK-Agents | Google ADK multi-agent prototype with background workers, MCP tools, memory hooks, and eval scaffolding. |
| GLM-OCR contribution | Merged fix to a 6k+ star OCR project: layout config id2label support, safer missing-mapping behavior, and unit tests. |
| bloom | AI content-marketing SaaS: multi-format text generation, image generation, and content improvement. Built in 48 hours. Live demo. |
| poco-5g-toggle | Android Quick Settings tile + widget (Kotlin, Shizuku) that toggles 5G for the active SIM on Poco/HyperOS. |
Archive: llm-experiments collects my early RAG, agent, and chatbot prototypes (11 projects, history preserved). portfolio has three iterations of my portfolio site.
- Inference & models: vLLM, AWQ quantization, prefix caching, chunked prefill, Unsloth fine-tuning, Hugging Face
- Agents & LLM apps: Google ADK, MCP, LangChain/LangGraph, CrewAI, RAG, embeddings, ChromaDB, FAISS, pgvector
- Languages: Python, TypeScript, Kotlin, SQL
- Backend & infra: FastAPI, Node.js, Docker, PostgreSQL, Vitest, GitHub Actions
- Cloud & platforms: Azure OpenAI / AI Foundry, Google Vertex AI, AWS, ServiceNow
- Adding more labs to inference-engineering
- Publishing held-out eval results for RocketGuard-1B
- Small, useful fixes to open-source AI tooling

