Explore my key research implementations, neural speech models, and engineering projects. This page lists various tools and frameworks I have developed, focusing on phonetic-preserving Text-to-Speech (TTS), fast-converging GAN vocoders (like FC-HiFiGAN, Swar, and Vaachika), speaker diarization pipelines, and multilingual audio deepfake detection (ADD) datasets. Most of these projects are open-source and available on GitHub for collaboration.
Research implementations, speech synthesis models, and engineering projects.
Speech & Audio Processing
- FC-HiFiGAN & PPHiFi-TTSPhonetic preserved & fast converging audio synthesis models with batchwise normalization
- BAANI Punjabi Vocoder296M-Parameter neural vocoder for end-to-end Punjabi speech synthesis
- Swar VocoderLongformer and Speaker-Aware GAN Vocoders for low-resource Indic speech synthesis (Gujarati & Hindi)
- VaachikaGAN-Based Neural Vocoder tailored for Marathi Text-to-Speech synthesis
- MLADDC & GGMDDCMulti-lingual audio deepfake detection corpora and evaluation frameworks
Machine Learning & Computer Vision
- Face-Mask Detector (YOLOv5)Real-time deep learning model detecting improper mask usage or unmasked human faces
- Custom Object DetectionTransfer-learning-based real-time object detection architecture using YOLOv5, SSD, and RCNN
- NLP Language IdentificationNLP algorithm to accurately identify language from text inputs
Web Applications & Distributed Systems
- Automated ML Data PipelineAutomated data pipeline using AWS Elastic Beanstalk, PHP, and Git for exploratory analysis
- CarefinderB2C medicine delivery platform with price comparison engine across local pharmacies (ASP.NET MVC, MySQL)
- NFT Marketplace & BlockchainSolidity smart contracts for NFT marketplaces & study on blockchain edge-cloud resource allocation