NLP & Multilingual Language Understanding
Developing robust, multi-model architectures and representation learning frameworks for low-resource, code-mixed, and real-world linguistic datasets.
Representative publications
Related projects
A quick map of what I work on, representative papers/projects, and what I’m looking for in a PhD lab.
I study trustworthy machine learning and applied statistical modeling, with current research spanning Natural Language Processing (multilingual & code-mixed NLP), epidemiological public-health modeling, and predictive data science.
My research starts from a foundational principle: how do we combine rigorous mathematical statistics with modern deep learning to build models that are reliable, robust, and empirically sound? In high-stakes fields like public health and low-resource multilingual communication, a model cannot merely optimize for benchmark accuracy—it must generalize under distribution shifts and provide transparent inference.
This perspective guides my work across three interconnected areas. In Natural Language Processing, I design multi-model and ensemble frameworks to classify complex sentiment in low-resource and code-mixed settings (such as English-Bangla app reviews), addressing noisy informal text where single models often fail. In Public Health & Epidemiology, I apply biostatistical modeling and machine learning to analyze health indicators, disease trajectories, and demographic health surveys. In Statistical Machine Learning, I focus on uncertainty estimation, robust inference, and principled evaluation protocols to ensure models do not rely on spurious correlations.
Across all projects, my priority is empirical rigor and thoughtful evaluation. A model that achieves high cross-validation metrics but breaks down in real-world deployment is not practically useful. My goal is to build methods whose assumptions, limits, and predictive uncertainties are transparent and dependable.
Looking ahead, I am actively seeking Ph.D. positions in Machine Learning, AI, and Applied Statistics, particularly in research groups where rigorous methodology, open artifacts, and meaningful applications to healthcare and NLP are taken seriously.
Developing robust, multi-model architectures and representation learning frameworks for low-resource, code-mixed, and real-world linguistic datasets.
Representative publications
Related projects
Applying rigorous statistical modeling, disease surveillance frameworks, and machine learning to analyze epidemiological data and healthcare outcomes.
Representative publications
Related projects
Bridging mathematical probability theory and deep learning to ensure reliable uncertainty estimation, model interpretability, and robust empirical validation.
Representative publications
Related projects