ML engineer adept at LLM pretraining, post-training (sft, dpo, rlvr), and agentic workflows.
- llm.pth - Hackable implementations of Autoregressive models (Llama, mixtral, gemma, deepseek), Research papers (cope, yarn, mod, mome, mla) and techniques (sft, dpo, kto, ipo) in Pytorch.
- llama3.cuda - llama3.cuda is an implementation of Llama 3.1 in pure C/CUDA. Consists of Swiglu, RoPE, CSE, RMSNorm and GQA kernels.
- prime-lab-trainer - A CC skill that builds, validates, and submits RL training environments on Prime Intellect β from any HuggingFace dataset to a live GRPO training run, fully automated.
- AutoSynth - Automatically create synthetic data using SOTA techniques (Self Instruct, Magpie, Agent Instruct, Arena Learning, Genstruct, Instruction Synthesizer, Self-Curation) using your LLMs.
- LightAgents - A wrapper free Agents library with RAG, function calling, json mode, telemetry and multi-layer memory.
- John Snow Labs - Lead post-training (sft, on-policy distillation, preference optimization, rlvr) for the JSL-Med 4B/8B/30B/70B family, building the data curation, reward design, and training pipelines behind each release β all ranked No. 1 on the Open Medical LLM Leaderboard. Built the vision-language extraction stack for handwritten medical documents (SFT + field-aware asymmetric OPD + constrained decoding over a 12k-entity medical vocabulary). Own medical evaluation end-to-end: clinical decision support, clinical calculation, medical error detection, race-bias, oncology, entity extraction, and multilingual benchmarks. Medical-LLMs
- Danucore - ML engineer adept at LLM pretraining, post-training (sft, dpo, rlvr), and agentic workflows.
- llm.pth - Hackable implementations of Autoregressive models (Llama, mixtral, gemma, deepseek), Research papers (cope, yarn, mod, mome, mla) and techniques (sft, dpo, kto, ipo) in Pytorch.
- llama3.cuda - llama3.cuda is an implementation of Llama 3.1 in pure C/CUDA. Consists of Swiglu, RoPE, CSE, RMSNorm and GQA kernels.
- prime-lab-trainer - A CC skill that builds, validates, and submits RL training environments on Prime Intellect β from any HuggingFace dataset to a live GRPO training run, fully automated.
- AutoSynth - Automatically create synthetic data using SOTA techniques (Self Instruct, Magpie, Agent Instruct, Arena Learning, Genstruct, Instruction Synthesizer, Self-Curation) using your LLMs.
- LightAgents - A wrapper free Agents library with RAG, function calling, json mode, telemetry and multi-layer memory.
- John Snow Labs - Lead post-training (sft, on-policy distillation, preference optimization, rlvr) for the JSL-Med 4B/8B/30B/70B family, building the data curation, reward design, and training pipelines behind each release β all ranked No. 1 on the Open Medical LLM Leaderboard. Built the vision-language extraction stack for handwritten medical documents (SFT + field-aware asymmetric OPD + constrained decoding over a 12k-entity medical vocabulary). Own medical evaluation end-to-end: clinical decision support, clinical calculation, medical error detection, race-bias, oncology, entity extraction, and multilingual benchmarks. Medical-LLMs
- Danucore - Architected and deployed a self-optimizing multimodal AI pipeline integrating RAG, agentic workflows, and open-source foundation models, leveraging LLM-as-a-Judge and Mixture-of-Agents for adaptive reasoning. Orchestrated 30+ GPUs across multi-node infrastructure to power end-to-end inference with Llama-3.1-70B, Phi-3-Medium-128K-Instruct, LLaVA-Next-8B, and SDXL-Lightning.
- QueryLoopAi - Pre-trained a 500M SLM from scratch on a carefully curated high-quality 15B tokens synthetic dataset. Created the entire training and evaluation pipeline along with managing training on 8xA100s. Created Kendrick, a mixture of experts model with 32k experts and Multi-latent head attention.
- Clinical Entity Extraction with Post-training and Constrained Decoding
- Post-training Qwen3.5-4B for Spatial Reasoning
- MHA vs MQA vs GQA vs MLA
- Linear Rope vs NTK vs YaRN vs CoPE
- MoE vs Dense vs Hybrid LLM architectures
- Multi-GPU Training of 70B LLM with Deepspeed and FSDP+Qlora
- Coding Deepseek-V2 from Scratch in PyTorch
- Best LLM Inference Engine? TensorRT vs vLLM vs LMDeploy vs MLC-LLM
View the archives (42 posts) @ zain.com.
- QueryLoopAi - Pre-trained a 500M SLM from scratch on a carefully curated high-quality 15B tokens synthetic dataset. Created the entire training and evaluation pipeline along with managing training on 8xA100s. Created Kendrick, a mixture of experts model with 32k experts and Multi-latent head attention.
- Clinical Entity Extraction with Post-training and Constrained Decoding
- Post-training Qwen3.5-4B for Spatial Reasoning
- MHA vs MQA vs GQA vs MLA
- Linear Rope vs NTK vs YaRN vs CoPE
- MoE vs Dense vs Hybrid LLM architectures
- Multi-GPU Training of 70B LLM with Deepspeed and FSDP+Qlora
- Coding Deepseek-V2 from Scratch in PyTorch
- Best LLM Inference Engine? TensorRT vs vLLM vs LMDeploy vs MLC-LLM
View the archives (42 posts) @ zain.com.
