

AI can get an answer right and still fail in the real world. Researchers at Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) in the UAE are witnessing this problem. While judging the candidate in the National AI Olympiad in Kazakhstan, Daniil Orel noticed that some parts of the code weren’t written by a human, but by instructions generated by an AI system.
In his statement, he mentioned, “At one point, I realized that as a human judge I could no longer be completely certain whether I was evaluating a student's own work or code generated with the assistance of a language model.”
This raised questions about how to check such work. Along with Orel, two other PhD students at Mohamed bin Zayed University of Artificial Intelligence are also pursuing the impacts of this similar problem. Apart from Orel, Amna Alhammadi, an Emirati, is shifting from machine learning to human-computer interaction to question whether technically capable AI systems are actually useful, understandable, and appropriate for the people they serve.
The third one, Emirati researcher Ali Aljaberi, is also examining vulnerabilities introduced through AI-assisted software development.
Also Read: Third-Party Cybersecurity Risks Rise as Digital Dependencies Grow
Orel’s AICD Bench tested AI-code detection using two million examples, 77 models, and nine programming languages. The results showed that detection tools had trouble when the language or topic changed. Code edited by people or made by both humans and AI was even harder to spot.
The bigger concern is security. Software can work properly and still have a weakness. A hidden flaw may not appear during a simple test. It could only become clear after the software is used by real people. Researchers say companies need to check AI-written code for security before putting it into use. Better coding results alone are not enough.
AI coding tools are becoming common in workplaces. As more developers use them, security checks will become even more important. The real test for AI is no longer just whether it can produce the right answer. It also needs to show that its work is safe when it leaves the testing stage and reaches real users.