Undergraduate Researcher

Decoding sequences across language, biology & security.

Natural Language Processing, Bioinformatics & Artificial Intelligence

Department of Computer Science and Engineering & Center for Advanced Bioinformatics and Artificial Intelligence Research (CABAIR), Islamic University, Kushtia, Bangladesh.

Portrait of Md. Nasit Jakoan
CABAIR · Islamic University
About

Resourceful by necessity, ambitious by design.

I am an undergraduate researcher focused on building innovative, high-impact AI systems. I specialize in solving complex problems using resourceful computational strategies, frequently leveraging accessible cloud environments like Kaggle and Google Colab. My work spans low-resource language NLP, computational genomics, and cybersecurity, with a strong drive toward publishing in high-impact Q1 journals. Currently, I am expanding my technical focus to include image recognition and medical imaging.

Research Focus

Four sequences, one method.

Whether the input is a sentence, a peptide chain, or a network session log, the underlying question is the same: what does this sequence mean, and can a model explain why.

01 — TEXT

Low-Resource NLP

Robust datasets and benchmarked Transformers (BanglaBERT, XLM-R) for Bengali dialect, political statement, hate speech, and emotion classification.

02 — PROTEIN

Bioinformatics & Genomics

Protein Language Models (ESM-2) and Explainable AI to decode antimicrobial peptides and resistance mechanisms in pathogens like E. coli.

03 — NETWORK

Cybersecurity

Hybrid probabilistic graph systems (CAPGBN) combining cross-attention and Bayesian networks for cloud anomaly detection.

04 — PIXEL

Medical Imaging & Vision

Deep neural networks and transformer architectures applied to medical image classification, including retinal disease detection.

Publications & Projects

Selected work.

Five case studies spanning corpus creation, model benchmarking, and explainability — each framed by the problem it addresses, the method used, and the measured outcome.

RBEC: Emotion Detection in Resource-Constrained Bengali Language

8,076 examples
The Problem

Bengali emotion detection lacked a clean, manually-verified corpus spanning the language's core emotional categories.

The Methodology

Built the Robust Bengali Emotion Corpus (RBEC) and benchmarked multiple model families against it for six-way emotion classification.

The Impact

Fine-tuned XLM-R reached state-of-the-art accuracy with strong agreement scores, setting a new benchmark for the task.

Fine-tuned XLM-R: 84.83% accuracy / F1, high Cohen's Kappa & MCC
View Publication

A Multi-Architecture System for Automated Crop Disease Identification

4,500 entries
The Problem

Farmers describe crop ailments in colloquial Bengali text, a format poorly served by existing agricultural diagnosis tools.

The Methodology

Curated a custom dataset of colloquial disease narratives and evaluated a multi-architecture classification system across it.

The Impact

Demonstrated that disease identification from informal, real-world farmer language is tractable at scale.

Custom dataset, multi-architecture benchmark across colloquial Bengali narratives
View Publication

A Comprehensive Framework for Bengali Religious Hate Speech

16,657 statements
The Problem

Religious hate speech in Bengali is widespread online but under-resourced for automated, reliable detection.

The Methodology

Developed a novel corpus of real-life statements and fine-tuned Bangla-BERT for binary hate speech classification.

The Impact

Achieved near-ceiling classification performance, offering a strong foundation for moderation tooling.

Fine-tuned Bangla-BERT: 98.89% accuracy / F1
View Publication

BanglaPoliText Benchmark

9,289 samples · 20 models
The Problem

No standardized benchmark existed for binary political statement classification in Bengali text.

The Methodology

Curated a dedicated dataset and evaluated 20 models, from traditional ML to Transformer-based systems, head-to-head.

The Impact

Established a reusable benchmark and clarified the performance gap between classical and Transformer approaches for this task.

20 models evaluated, spanning classical ML to Transformers
View Publication

Security Risk Assessment via CAPGBN

5,000 sessions
The Problem

Cloud network anomaly detection must contend with severely imbalanced traffic data and opaque model decisions.

The Methodology

Proposed CAPGBN, a Cross-Attention-Based Probabilistic Graph and Bayesian Network framework for risk assessment.

The Impact

Validated the hybrid approach against a highly imbalanced dataset of real network session records.

Evaluated on 5,000 highly imbalanced network session records
Arsenal

Tools of the trade.

Machine Learning & AI

NLP Image Recognition LLMs Protein Language Models (ESM-2) CNN BiLSTM GRU GCN Transformers

Tools & Frameworks

Python PyTorch TensorFlow Scikit-Learn Git

Domain Expertise

Statement Classification Emotion Detection Dataset Curation Sequence Classification Explainable AI