Backend • Edge Computing • Distributed Systems

Building backend systems for the edge and beyond.

I'm Mohammad Sufiyan ,an Edge AI/ML engineer and backend developer. I love building backend systems, with a particular focus on edge computing and distributed systems. I'm especially interested in designing software that runs reliably on resource-constrained devices and scales across distributed environments.

01

Experience

Backend Developer Intern
Rudra Nursery
June 2025 — July 2025
  • Independently architected and deployed a full-stack e-commerce platform for a live nursery business using Django, handling auth, payments, media, and deployment end to end.
  • Implemented secure sign-in with Django AllAuth + Google OAuth and integrated Razorpay for payment processing.
  • Set up Neon serverless PostgreSQL for the database, Cloudinary for media, and Docker for containerized deployment on Render.
02

Projects

Infrastructure for running and scheduling LLM inference on hardware that was never meant to hold it.

ModelPulse
light weight Edge inference transfer library
A Python library that automates pushing GGUF models from a server to a resource-constrained edge client, running inference via llama.cpp, and returning structured telemetry - no manual steps between model swaps.
llama.cppFastAPI / WebSocketstmpfsQuantization
  • Zero-disk inference path: shards stream over HTTP, assemble in RAM via Linux /dev/shm, protecting edge storage from write cycles.
  • Tensor/block-level delta updates - changed shards are retransmitted after quantization, not the full model.
  • WebSocket control plane for model-ready broadcasts and live, in-memory model swaps with no process restart.
ModelFlow
RL environment for inference scheduling
A custom OpenAI Gym – compatible environment (built with OpenEnv) where an agent learns to load, execute, evict, or replace LLMs to maximize throughput under real CPU/RAM constraints — without hitting OOM.
Reinforcement LearningDockerFastAPI
  • Discrete action space (load / execute / evict / replace / idle) with observations reflecting memory pressure and queue depth.
  • Programmatic graders scoring agents on quantization trade-offs, OOM avoidance, and queue efficiency.
  • Benchmarked on an Intel i3 / 8GB laptop so constraints mirror real edge conditions, not lab settings.
MEDICO-AGENT
Medical VQA · vision-language fine-tuning
A multi-modal medical assistant answering queries from text, X-ray/CT/MRI images, or both, via an agent-based architecture over fine-tuned Qwen-VL models.
LoRA / UnslothRAGGradio
  • Fine-tuned four Qwen-VL models on 12,149 medical VQA samples plus a 6,666-sample PubMed Vision subset.
  • Best model (Qwen3-VL-8B Instruct) reached 45.90% exact match and 94.79% average BERTScore.
  • RAG pipeline for lung-disease grounding via Tavily search, observed with LangSmith; FastAPI + Gradio for inference.
More on github.com/MdSufiyan005 - repos including hackathon builds and smaller tools.
03

Skills & background

Languages
PythonSQLJavaScriptC++
ML / AI
PyTorchTransformersHugging FaceLoRAUnslothGGUFllama.cppReinforcement Learning
Inference & systems
QuantizationEdge inferenceModel shardingtmpfs / /dev/shmWebSocket orchestration
Backend & DevOps
FastAPIDjangoDockerRedisCeleryUvicorn
Databases
PostgreSQLMongoDBMySQL

Education

B.Tech, Computer Science & Engineering (AI & Edge Computing)
MIT ADT University, Pune — CGPA 8.65/10
2023 — 2027

Certifications

Achievements

  • Built a vendor management platform at the TuteDude Online Website Hackathon.
  • Developed a real-time exam proctoring system at The Great Bengaluru Hackathon.
  • Led a team at Smart India Hackathon (SIH) to design a ship navigation & safety assistance system.