Research implementations, speech synthesis models, and engineering projects.
Speech & Audio Processing
FC-HiFiGAN & PPHiFi-TTS
Phonetic preserved & fast converging audio synthesis models with batchwise normalization
BAANI Punjabi Vocoder
296M-Parameter neural vocoder for end-to-end Punjabi speech synthesis
Swar Vocoder
Longformer and Speaker-Aware GAN Vocoders for low-resource Indic speech synthesis (Gujarati & Hindi)
Vaachika
GAN-Based Neural Vocoder tailored for Marathi Text-to-Speech synthesis
MLADDC & GGMDDC
Multi-lingual audio deepfake detection corpora and evaluation frameworks
Machine Learning & Computer Vision
Face-Mask Detector (YOLOv5)
Real-time deep learning model detecting improper mask usage or unmasked human faces
Custom Object Detection
Transfer-learning-based real-time object detection architecture using YOLOv5, SSD, and RCNN
NLP Language Identification
NLP algorithm to accurately identify language from text inputs
Web Applications & Distributed Systems
Automated ML Data Pipeline
Automated data pipeline using AWS Elastic Beanstalk, PHP, and Git for exploratory analysis
Carefinder
B2C medicine delivery platform with price comparison engine across local pharmacies (ASP.NET MVC, MySQL)
NFT Marketplace & Blockchain
Solidity smart contracts for NFT marketplaces & study on blockchain edge-cloud resource allocation