AI in CS Education · University of Dhaka

Md Fahim Arefin

I study how generative AI changes the learning, teaching, and assessment of computer science.

My current work combines cognitive diagnostic modeling, programming-submission traces, and language-model evaluation to distinguish a correct answer from demonstrated understanding. Recent work appears in ACM TOSEM, Findings of EMNLP, TMLR, and workshops at ICSE, ASE, and EDM.

Md Fahim Arefin
Md Fahim Arefin
Assistant Professor, Computer Science & Engineering
21peer-reviewed publications
100+Google Scholar citations
3funded research grants

About

I am an Assistant Professor of Computer Science & Engineering at the University of Dhaka. I study how AI is changing CS education: how we diagnose programming skills, learn from students’ attempts, evaluate AI tutors, and grade work without confusing plausible output with understanding.

Most of my publications grow from undergraduate thesis supervision and student-led research. I also collaborate with faculty across universities in Bangladesh and abroad, and I am interested in expanding that network through joint studies and PhD research.

Outside the university, I co-founded AlterYouth, which supports 2,500+ monthly scholarships for underprivileged primary school students in Bangladesh.

How reliable is AI as a tutor and as an evaluator?

I attended the Festival of Learning earlier this year, and since then I’ve been focused on this question: how can computing education measure and support genuine learning when AI can generate plausible code, explanations, and feedback on demand?

CodeCognify: Can small language models diagnose programming skills?

I am comparing DINA and G-DINA with three small language models and two large models on response data from 93 students, 49 problems, and a 12-skill Q-matrix. The aim is to test whether more accessible models can approximate fine-grained psychometric diagnosis.

Verifiable grading of handwritten CS exams

I am benchmarking multimodal LLMs on 54 anonymised handwritten scripts and developing EDU, a three-stage architecture that separates layout segmentation, visual-syntax recovery, and verification. It is designed to execute checkable code and abstain when recognition is uncertain.

Why Didn’t It Pass?

I am reconstructing pre-acceptance trajectories from 6,491 submissions by 852 students across 60 programming problems. Instead of treating wrong answers as noise, the project studies error categories, persistence, and the point at which debugging shifts from patching to conceptual revision.

Failure-aware evaluation of LLM tutors

I am evaluating whether AI programming tutors detect, repair, or reinforce the misconception behind an incorrect solution. The benchmark draws from 3,101 annotated submissions and scores 2,350 tutor responses across three models and four prompting strategies.

Selected work

These five recent papers show the range of collaborations and publication venues underpinning the focused research agenda above.

The complete publication record remains available here and on Google Scholar.

BanglaMemeX: Advancing Cultural Metaphoric Image Interpretation in Bangla with a Multimodal Explainable Dataset

Findings of EMNLP 2026 · To appear

Sakib, Mahmud, Arefin, Fahim

A 3,000-item explainable dataset for culturally grounded interpretation of Bangla memes.

Strict Graders and Hallucinated Mastery: A Cognitive Diagnostic Study of Language Models on Dynamic Programming

CSEDM at EDM 2026 · To appear

Solaiman, Arefin

Shows language-model graders dividing between excessive strictness and crediting skills that submissions never demonstrate.

Full publication record

UHGMiner: A Framework and Algorithm for High-Utility Hypergraph MiningADMA 2026 · Jawad, Habib, Alam, Arefin, Ahmed, Leung

BanglaMemeX: Advancing Cultural Metaphoric Image Interpretation in Bangla with a Multimodal Explainable DatasetFindings of EMNLP 2026 · Sakib, Mahmud, Arefin, Fahim

VibeCheck: Assessing the Quality of LLM-Generated Unit TestsPOVC at ASE 2026 · Tabassum, Intesum, Arefin, Zaman

Strict Graders and Hallucinated MasteryCSEDM at EDM 2026 · Solaiman, Arefin

Hyperedge Anomaly Detection with Hypergraph Neural NetworkTransactions on Machine Learning Research · Alam, Rahman, Arefin, Ahmed, Mahmud, Sakib, Leung

A-ProS: Towards Reliable Autonomous Programming Through Multi-Model FeedbackACM TOSEM · Tabassum, Hossain, Arefin, Islam, Zaman

LLM-ProS: Analyzing Large Language Models’ Performance in Competitive Problem SolvingLLM4Code at ICSE 2025 · Hossain, Tabassum, Arefin, Zaman

A New Tree-Based Approach to Mine Sequential PatternsExpert Systems with Applications · Rizvee, Ahmed, Arefin, Leung

Unveiling Market Sentiments: Stock-Market Responses to Diverse News EventsIMCOM 2024 · Solaiman, Rahman, Arefin, Ahmed, Leung, Madill

Applying Sequential Correlation to Identify Order Dependent Activity PairsDhaka University Journal of Applied Science and Engineering · Arefin, Islam, Ahmed

Mining Contextual Item Similarity without Concept HierarchyIMCOM 2022 · Arefin, Ahmed, Rizvee, Leung, Cao

Comparison of Different Sentiment-Analysis Techniques for Bangla ReviewsIEEE R10 HTC 2022 · Jabin, Suhi, Arefin, Hasib

An Automated System to Calculate Marks from Answer ScriptsIEEE IEMCON 2022 · Rizvee, Arefin, Khan, Islam, Rabbi

Data-Driven Hybrid Optimisation-Based Deep Network for Short-Term Residential Load ForecastingIEEE IEMCON 2021 · Sakib, Hasib, Tasawar, Tanzeem, Arefin, Islam

Tree-Miner: Mining Sequential Patterns from SP-TreePAKDD 2020 · Rizvee, Arefin, Ahmed

Mining Sequential Correlation with a New MeasureICDM 2018 · Arefin, Islam, Ahmed

Google Scholar record ↗

Research grows through supervision and collaboration.

Most of my publications have grown from undergraduate thesis supervision and student-led research. I treat supervision as collaborative research training: students learn to frame a question, build defensible evidence, and carry the work through publication.

I have also collaborated with faculty across universities in Bangladesh and abroad. Those collaborations connect educational data mining, software engineering, language-model evaluation, pattern mining, and culturally grounded AI.

I am looking to expand this network around a more focused agenda: assessment beyond correctness, learning analytics for programming, reliable AI feedback, and the changing role of the instructor when generative models are part of everyday student work.

Discuss a research idea ↗

Academic record

Both degrees are from the University of Dhaka, where I now teach. My research has been supported by university, government, and national ICT programmes.

Education and recognition

M.Sc. in Computer Science & EngineeringUniversity of Dhaka

B.Sc. in Computer Science & EngineeringUniversity of Dhaka · Gold Medal for Academic Excellence

Dean’s AwardFaculty of Engineering and Technology

Research grants

Consultant · RAG-driven business intelligenceICSETEP, University Grants Commission

Domain Expert · Enhancement of Bangla Language in ICTBangladesh Computer Council

Co-Principal Investigator · Utility-based hypergraph miningUGC Research Grant

Let’s compare research questions.

If your work touches AI in computing education, programming-skill assessment, learning analytics, or reliable model evaluation, I would be glad to hear from you.

In Europe in October 2026: ASE in Munich, 12–16 October, and EMNLP in Budapest, 24–29 October. If you will be at either, let’s meet for coffee.