Sumit Vaise
Senior ML Engineer with 9+ years building end-to-end AI/ML systems. Started with deep learning and computer vision (CNNs, Siamese Networks, EfficientNet) at Bosch and LNT, through to production RAG pipelines and LLM applications at Fidelity. I've trained models for medical imaging, autonomous driving, insurance, and fintech, then built the MLOps infrastructure to keep them running at scale on GCP and AWS.
I've led teams of up to 12, mentored engineers, delivered 6 production ML systems across verticals, and run ML guild sessions on LLM fine-tuning and retrieval. Currently based in Scarborough, Ontario. Canadian PR. Open to Sr ML / AI Engineer roles across Canada and the US, remote preferred.
Where I've worked
- Built a Multi-Document RAG Pipeline using LangChain, Sentence Transformers, and ChromaDB, cutting analyst research time by 40% across 500+ financial reports.
- Deployed a Vector ETL pipeline for multi-source document ingestion, implementing recursive text chunking and automated embedding generation.
- Developed end-to-end ETL and ML pipelines for real-time communication error detection using AWS, Airflow, Snowflake, Flask, and AWS EKS.
- Established MLOps best practices: MLflow experiment tracking, GitHub Actions CI/CD, automated retraining triggers, Grafana monitoring dashboards.
- Designed and maintained scalable ML pipelines on GCP (BigQuery, Vertex AI, Docker, Kubernetes, KubeFlow) serving 14M+ records.
- Implemented end-to-end Random Forest pipeline predicting insurance policy lapses, identifying the top 2% at-risk customers and directly impacting business retention strategy.
- Designed comprehensive monitoring, alerting, and observability systems for ML model health, maintaining strict production SLAs.
- Partnered with product managers to align ML technical requirements with business stakeholder needs in Agile sprints.
- Inventory Allocation Prediction (Team Lead, Team of 4): Architected and led end-to-end ML solution for US shipment optimisation, translating business requirements into production systems.
- Insurance Claims Processing (Team of 12): Implemented a scalable Deep Learning classification system on GCP using TensorFlow, OCR, and PyTorch, achieving a 70% speed-up in claims processing.
- Mentored 3 junior engineers; led bi-weekly ML guild sessions on LLM fine-tuning and retrieval-augmented generation.
- Delivered 6 production ML systems across healthcare, fintech, and media verticals on GCP and AWS.
- Built a computer vision pipeline for hairstyle classification and recommendation using PyTorch CNNs; integrated with iOS/Android app serving 50K+ users.
- Quantized models with TensorFlow Lite for mobile deployment, reducing model size while preserving accuracy.
- Built and trained a person re-identification and tracking system using a CNN-based Siamese Network (PyTorch) with a custom Qt/C++ GUI for real-time inferencing.
- Reduced object detection model training time by 22% without accuracy loss through dataset optimisation (dense-to-sparse) on AWS EC2.
- Led projects in computer vision: medical image stitching (MATLAB, OpenCV, C++) and thermal person detection on Raspberry Pi at 82% accuracy.
- Developed 3-class object detection using transfer learning (VGG16) on NVIDIA Jetson-TK1 achieving 96% accuracy.