CodeCognify: Can small language models diagnose programming skills?
I am comparing DINA and G-DINA with three small language models and two large models on response data from 93 students, 49 problems, and a 12-skill Q-matrix. The aim is to test whether more accessible models can approximate fine-grained psychometric diagnosis.