Hi, my name is

Mandeep Jaglan.

AI & ML Engineer — video pipelines, computer vision & MLOps.

AI & ML engineer with 4+ years of experience across data analytics, NLP pipelines, computer vision systems, and MLOps infrastructure. I build end-to-end video processing pipelines, RAG-based SQL intent systems, OCR and ReID workflows, and scalable ETL infrastructure on GCP. Published author of two open-source Python packages on PyPI.

01. Work Experience

AI Infrastructure Engineer · Streamingo.ai

Oct 2024 — Present

Video Processing & ML Infrastructure

  • Built a RabbitMQ + Redis based cron pipeline ("fizzing") to orchestrate VLM inference on 2,000–4,000 videos per day, generating structured insights and result sheets.
  • Reduced end-to-end pipeline runtime from ~1.5 days of manual processing to under 4 hours via fully automated parallel execution across 7–8 GPUs on 2–3 machines simultaneously.
  • Developed GCS video management scripts for merging, cutting, and preprocessing timestamped video segments prior to inference.
  • Built an end-to-end Jenkins CI/CD pipeline — from video recording through VLM inference to results delivery — with step-wise email notifications and queue polling.
  • Engineered supporting Jenkins jobs: abort/delete stuck queue items, GCS upload/download, dynamic YML generation for Docker startup based on machine type and available GPUs, and data upload to an analytics dashboard.

Basketball Minimap Pipeline (Computer Vision)

  • Designed and built a multi-stage CV pipeline transforming basketball broadcast video into top-down 2D minimap overlays using YOLOv9/v11, SAM 2, BoxMOT, and homography transformation — from unusable prototype to production-ready pilot.
  • Implemented Global Track Association (GTA) to maintain consistent player identities across scenes and occlusions.
  • Built team ReID using KMeansConstrained on HSV histograms; jersey OCR pipeline using FSRCNN super-resolution cross-validated against substitution logs.
  • Applied bidirectional Kalman filtering on court keypoints and linear interpolation of player tracks (up to 45-frame gaps) for smooth minimap rendering.
  • Produced analytics output: shot/turnover location maps, assist arrows, player marking analysis, Gaussian heatmaps, and per-player speed/distance tracking.
  • Integrated a DeepLabel review loop for high-accuracy labeling; designed a SAM 2.1 targeted labeling system to reduce manual effort on future iterations.

NLP & Conversation Analytics

  • Built an NLP → intent → BigQuery SQL pipeline for querying fizzing result sheets, with schema-aware prompt engineering, pre/post-processing, and automated chart and insight generation.

SDE 2 – Machine Learning · Nanonets (via Crewscale)

Mar 2024 — Sep 2024
  • Improved OCR accuracy to 95%+ across 10+ client datasets by optimising LLM configuration, prompt engineering, and building Python pre/post-processing scripts.
  • Fine-tuned a GPT-3 based model for named entity recognition, increasing name identification accuracy by 35%.
  • Automated data workflows using the Nanonets API and AWS S3 to handle 10 TB+ of data efficiently.
  • Built an NLP-driven recruitment search tool (beanbag.ai) converting natural language queries to Elasticsearch and SQL via the Claude API, with PDL person API integration.

Associate Data Analyst · Comscore

Jan 2022 — Feb 2024
  • Built a Facebook Ads data retrieval solution using APIs and web scraping, reducing processing time by 86%.
  • Wrote scripts to collect and filter raw panel data from Hadoop, enabling access to 36+ TB of unstructured data.
  • Automated data analysis and reporting with Python, SQL, and Snowflake; created and debugged 100+ SSRS reports.
  • Developed a PySpark logistic regression model predicting panelist churn, improving prediction accuracy by 25%.
  • Built an ML clustering solution to categorize users by web activity, informing UX and marketing strategies.

02. Projects

Open Source · PyPI

VideoProcToolkit

Python library for video download, multi-backend frame extraction (FFmpeg/OpenCV/Decord), trimming, concatenation, FPS/resolution conversion, thumbnails, and metadata.

View on PyPI →

Open Source · PyPI

basketball_datakit

NBA/NCAA data toolkit: play-by-play, box scores, event classification, lineup tracking, ESPN scraper fallback, and name normalization. (v0.2.3)

View on PyPI →

NLP · LLM

Recruitment Search Automation — beanbag.ai

NLP query-to-Elasticsearch/SQL pipeline for recruiter automation, integrated with the PDL person API via the Claude API.

Machine Learning

Panelist Churn Prediction — Comscore

PySpark logistic regression model improving churn prediction accuracy by 25%, built on Hadoop and Snowflake.

Machine Learning

IPL Winner Prediction

Random Forest model (80% accuracy) for IPL match outcome prediction with full data analysis and evaluation.

View on GitHub →

03. Skills

Languages

PythonSQLPySparkC/C++Shell Scripting

Data & Cloud

BigQueryGCP/GCSAWS S3SnowflakeHadoopHiveDocker

MLOps & Infra

Jenkins CI/CDRabbitMQRedisETL Pipelines

CV & ML

YOLOv9/v11SAM 2BoxMOTHomographyOCRReIDKalman Filter

NLP & LLM

RAG PipelinesPrompt EngineeringLLM Fine-tuningClaude API

Libraries & Tools

PandasNumPyScikit-learnTensorFlowFFmpegOpenCVGitLinuxPower BISSRS

04. Education & Certifications

B.Tech — Computer Science & Engineering

IIIT Bhagalpur · 2018 – 2022

Certifications

  • SQL for Data Science — Coursera
  • Algorithmic Toolbox — Coursera
  • GCP Cloud Engineering & ML — Qwiklabs

05. Beyond Work

  • Mentored 5 interns in data validation and quality control
  • Coordinated Google DSC AI/ML Club
  • Led team at Smart India Hackathon (Android app)
  • Managed programming community for GeeksforGeeks events
  • NSS volunteer — education for underprivileged children