01. Work Experience
AI Infrastructure Engineer · Streamingo.ai
Oct 2024 — PresentVideo Processing & ML Infrastructure
- Built a RabbitMQ + Redis based cron pipeline ("fizzing") to orchestrate VLM inference on 2,000–4,000 videos per day, generating structured insights and result sheets.
- Reduced end-to-end pipeline runtime from ~1.5 days of manual processing to under 4 hours via fully automated parallel execution across 7–8 GPUs on 2–3 machines simultaneously.
- Developed GCS video management scripts for merging, cutting, and preprocessing timestamped video segments prior to inference.
- Built an end-to-end Jenkins CI/CD pipeline — from video recording through VLM inference to results delivery — with step-wise email notifications and queue polling.
- Engineered supporting Jenkins jobs: abort/delete stuck queue items, GCS upload/download, dynamic YML generation for Docker startup based on machine type and available GPUs, and data upload to an analytics dashboard.
Basketball Minimap Pipeline (Computer Vision)
- Designed and built a multi-stage CV pipeline transforming basketball broadcast video into top-down 2D minimap overlays using YOLOv9/v11, SAM 2, BoxMOT, and homography transformation — from unusable prototype to production-ready pilot.
- Implemented Global Track Association (GTA) to maintain consistent player identities across scenes and occlusions.
- Built team ReID using KMeansConstrained on HSV histograms; jersey OCR pipeline using FSRCNN super-resolution cross-validated against substitution logs.
- Applied bidirectional Kalman filtering on court keypoints and linear interpolation of player tracks (up to 45-frame gaps) for smooth minimap rendering.
- Produced analytics output: shot/turnover location maps, assist arrows, player marking analysis, Gaussian heatmaps, and per-player speed/distance tracking.
- Integrated a DeepLabel review loop for high-accuracy labeling; designed a SAM 2.1 targeted labeling system to reduce manual effort on future iterations.
NLP & Conversation Analytics
- Built an NLP → intent → BigQuery SQL pipeline for querying fizzing result sheets, with schema-aware prompt engineering, pre/post-processing, and automated chart and insight generation.
SDE 2 – Machine Learning · Nanonets (via Crewscale)
Mar 2024 — Sep 2024- Improved OCR accuracy to 95%+ across 10+ client datasets by optimising LLM configuration, prompt engineering, and building Python pre/post-processing scripts.
- Fine-tuned a GPT-3 based model for named entity recognition, increasing name identification accuracy by 35%.
- Automated data workflows using the Nanonets API and AWS S3 to handle 10 TB+ of data efficiently.
- Built an NLP-driven recruitment search tool (beanbag.ai) converting natural language queries to Elasticsearch and SQL via the Claude API, with PDL person API integration.
Associate Data Analyst · Comscore
Jan 2022 — Feb 2024- Built a Facebook Ads data retrieval solution using APIs and web scraping, reducing processing time by 86%.
- Wrote scripts to collect and filter raw panel data from Hadoop, enabling access to 36+ TB of unstructured data.
- Automated data analysis and reporting with Python, SQL, and Snowflake; created and debugged 100+ SSRS reports.
- Developed a PySpark logistic regression model predicting panelist churn, improving prediction accuracy by 25%.
- Built an ML clustering solution to categorize users by web activity, informing UX and marketing strategies.
02. Projects
Open Source · PyPI
VideoProcToolkit
Python library for video download, multi-backend frame extraction (FFmpeg/OpenCV/Decord), trimming, concatenation, FPS/resolution conversion, thumbnails, and metadata.
View on PyPI →Open Source · PyPI
basketball_datakit
NBA/NCAA data toolkit: play-by-play, box scores, event classification, lineup tracking, ESPN scraper fallback, and name normalization. (v0.2.3)
View on PyPI →NLP · LLM
Recruitment Search Automation — beanbag.ai
NLP query-to-Elasticsearch/SQL pipeline for recruiter automation, integrated with the PDL person API via the Claude API.
Machine Learning
Panelist Churn Prediction — Comscore
PySpark logistic regression model improving churn prediction accuracy by 25%, built on Hadoop and Snowflake.
Machine Learning
IPL Winner Prediction
Random Forest model (80% accuracy) for IPL match outcome prediction with full data analysis and evaluation.
View on GitHub →03. Skills
Languages
Data & Cloud
MLOps & Infra
CV & ML
NLP & LLM
Libraries & Tools
04. Education & Certifications
B.Tech — Computer Science & Engineering
IIIT Bhagalpur · 2018 – 2022
Certifications
- SQL for Data Science — Coursera
- Algorithmic Toolbox — Coursera
- GCP Cloud Engineering & ML — Qwiklabs
05. Beyond Work
- Mentored 5 interns in data validation and quality control
- Coordinated Google DSC AI/ML Club
- Led team at Smart India Hackathon (Android app)
- Managed programming community for GeeksforGeeks events
- NSS volunteer — education for underprivileged children