SDE II · Hyderabad, India

Mohammed
Zafeeruddin

Software engineer · Platform / DevOps / MLOps

I write the Python services, pipelines and Kubernetes clusters behind production AI systems. Right now that means a computer-vision platform running on 50+ GPU nodes and 2,000+ cameras, plus air-gapped deployments for Saudi government clients.

zafeer@nabeh: ~
$ kubectl get engineer zafeer -o yaml
apiVersion: people/v2
kind: Engineer
metadata:
  name: mohammed-zafeeruddin
  labels:
    role: sde-2
    focus: platform.devops.mlops
spec:
  product: baseer-builder
  writes: [python, bash, typescript]
  runs: [k8s, argocd, vault, ansible]
  replicas: 1  # very caffeinated
status:
  phase: Running
  since: 2024-05
$ 
50+k8s nodes operated
2,000+cameras on the platform
55h → 5hmodel training with DDP
3 days → 20 minfull platform setup
SAR 300Ksaved per year at Expro
$ tree baseer-builder/ --owned-by=me

The product I build

Baseer Builder

A no-code computer-vision platform: collect data, label it, train models, run custom notebooks, deploy inference to live cameras and look at the results on dashboards. I own the MLOps, DevOps and infrastructure side, and I wrote the services below.

largest deploymentEPM: Eastern Province Municipality, KSA

A government project with 200+ use cases on 2,000+ cameras, sized for 50+ A100-class machines, delivered on Baseer Builder.

warehouse/training/k8s Job · on demand

Model Training

  • Training microservice for detection, classification and segmentation, run as a Kubernetes Job so it survives interruptions
  • PyTorch DDP across pooled GPUs cut a 55h run to 5h
  • Pause, stop, resume and prioritise; live metrics in MLflow; progress over Kafka via logger + status services
warehouse/notebooks/StatefulSet

Jupyter Notebooks

  • Each notebook gets its own subdomain through ingress + MetalLB on bare metal
  • Code persisted on NFS-backed NAS
  • Idle detection warns the user, then shuts the notebook down to free GPUs
deploy/devices/service + k8s Jobs + Ansible

Deploy Manager

  • Rewrote Bash onboarding as a deploy-manager service that runs Ansible from Kubernetes Jobs
  • Single and bulk onboarding, patching to a pinned version, and full cleanup on deboarding
  • 7-phase lifecycle reported to the backend over Kafka; device health from kube-prometheus-stack
deploy/cameras/persistent service

Camera Service

  • RTSP → HLS through MediaMTX with an API for the backend to add or remove streams
  • Tuned for thousands of cameras with parallel conversion
warehouse/faces/k8s Job · on demand

Face Registration

  • Supports multiple face models and writes embeddings to Qdrant
platform/infra/the foundation

Cluster & Delivery

  • HA kubeadm cluster (HAProxy) across dev, prod and EPM
  • Jenkins CI → image tags in an infra repo → Argo CD deploys
  • Harbor registry on NAS used by 8 teams; Vault for every secret
  • Moved InfluxDB to ClickHouse and every Docker-based service into k8s
$ git log --author=zafeer --oneline

Experience

Masterworks · NabehMay 2024 — present

An AI startup building computer-vision and GenAI products for enterprise and government clients across Saudi Arabia and the UAE.

  1. Apr 2026 — now
    SDE II

    I own the platform and infrastructure for Baseer Builder and its client deployments. That covers architecture, Kubernetes, CI/CD, secrets, networking, storage and on-call debugging.

  2. Mar 2025 — Mar 2026
    Software Engineer

    Converted to full time early. Built the training, notebook, camera and device services, the HA NGINX setup and the delivery pipelines, and moved the platform onto Kubernetes.

  3. May 2024 — Feb 2025
    Engineering Intern

    Built four computer-vision PoCs: crowd analytics for Ajdan, access and presence for King Salman Military Base, multi-camera footfall for Gold Chain, and turnaround monitoring at Dubai Airport.

Also on my plate

  • HA NGINX across machines using Keepalived virtual IPs, with config synced from GitHub
  • Office network moved to VLANs and made ISP-agnostic
  • SSH key-only access, per-team users and AD-backed machine login
  • A DNS registrar move that saves about $500/mo, and a self-hosted Harbor that saves about $1K/mo
  • A 3.5 TB ClickHouse and /serve migration to NVMe for Humain, with checksums and rollback copies
  • Multi-VPC ACK + ECS environments with GitLab CI for Monsha'at
$ cat postmortems/*.md | head

Bugs I've chased down

Saudi MEP · air-gapped

Recording bots could only ever run on one node

found
Every bot mounted the API's ReadWriteOnce volume, and the API pinned itself with a required podAffinity. The whole cluster's capacity was effectively one machine.
fixed
Bots now record to node-local ephemeral storage. The API streams finished recordings back over the k8s exec channel and unpacks the tar in flight. Bots spread across all 4 workers, and one env var switches back to the old layout.
Saudi MEP · air-gapped

A manifest that would have deleted 22 live secrets

found
The checked-in manifest had drifted from production and would have dropped every Graph, Teams, Webex and SQL credential reference.
fixed
Shipped a targeted strategic-merge patch instead, then reconciled the repo back to zero drift.
Expro · Alibaba + Azure DevOps

Deploys silently failing to read config

found
A periodic Vault token had expired with nothing renewing it. Separately, a config render emptied the env file whenever Vault was unreachable.
fixed
Automated token renewal, made renders fail closed, and made Vault the only source of config. A refresh pipeline recreates containers without a rebuild.
Expro · Alibaba + Azure DevOps

Office subnet blackholed by Docker

found
The default Docker bridge was handing out /16 ranges that overlapped the office network.
fixed
Moved Docker onto non-overlapping address pools. I also fixed deploy stages that ran after failed builds and recovered a volume orphaned by a compose rename.
$ ls ~/side-projects

Things I've built on my own

Causeway

open-sourcing soon
network · streaming

View and record cameras that sit behind VPNs, jump hosts and SSH tunnels. Each customer gets its own Linux network namespace, so a dropped tunnel can't leak traffic.

  • 8-gate diagnostics, from VPN to stream handshake
  • WebRTC with LL-HLS fallback
  • Crash-safe MPEG-TS segments → MP4 → S3
  • 298 automated tests
Python 3.12Next.js 16WireGuardWebRTCS3

kuiqctl

released
open source · CLI

Creates a persistent single-node kubeadm cluster that stays Ready when your laptop changes Wi-Fi or gets a new DHCP lease.

  • Stable loopback node identity + mDNS
  • systemd watcher reconciles addresses
  • Preflight gates before anything destructive
  • Versioned releases with CI
BashPythonkubeadmcontainerdCalico

CloudPad

built
dev platform

Chat to build a site and get a live URL plus a browser VS Code for it. Two pods share one volume: one serves the code, the other writes it.

  • Per-workspace live links over ingress
  • Edit by hand or let the model do it
  • Changes show up immediately
KubernetesFastAPIReactAWS EC2MetalLB

AWS EKS Doc Platform

built
cloud architecture

A private EKS platform for document processing. API Gateway → VPC Link → internal ALB, with workers pulling jobs from SQS.

  • Multi-AZ VPC with public, app and DB subnets
  • Pod Identity instead of static keys
  • Debugged node joins, CNI and ALB discovery
EKSRDSS3SQSIAM

LensCopy

built
linux desktop

A local image viewer for Linux. Click or drag over an image to copy the text in it, with OCR running fully offline.

  • OCR on a worker thread so the UI stays responsive
  • Installers for apt, dnf and pacman
  • Nothing leaves the machine
PythonGTK 3TesseractPillow
full stack

A blogging platform with live notifications, comments and replies, and Google + OTP sign-in, running on Cloudflare's edge.

  • Workers + Pages, images in R2
  • Postgres via Prisma Accelerate
  • CI/CD on GitHub Actions
ReactHonoCloudflareRedisPostgres
$ cat stack.toml

Tools I use

[languages]
PythonBashTypeScriptSQL
[orchestration]
KuberneteskubeadmHelmDockercontainerdKubeflow
[iac & config]
TerraformAnsibleHashiCorp Vault
[delivery]
JenkinsArgo CDGitHub ActionsAzure DevOpsGitLab CIHarbor
[networking]
NGINXHAProxyKeepalivedMetalLBVLANsWireGuard
[data & ml]
PyTorch DDPMLflowKafkaMQTTClickHousePostgresRedisQdrant
[observability]
PrometheusGrafanaAlertmanager
[cloud]
AWS (EKS, VPC, RDS, S3, SQS)Alibaba (ECS, ACK, ACR)Cloudflare
$ ./contact.sh

Working on something hard to run?

I'm happy to talk about platform, DevOps or MLOps roles, or about a cluster that keeps falling over.

mohammed.xafeer@gmail.com