AI is now the tool of choice for millions of people around the world. From writing emails to studying history to summarising lengthy documents, tools such as ChatGPT make answering questions seem like a breeze. But AI is not infallible; it might provide incorrect answers, alter its response if presented with the same question twice, or provide a slow answer when queries are more complex. To assess the reliability of ChatGPT, a team of researchers (Aladdin, Muhammed, Abdulla, and Rashid) created a testing framework dubbed Precision Answer Comparison and Evaluation Model (PACEM). Here's a quick summary of what they discovered.
Unlike traditional checking for matching words, PACEM evaluates answers based on semantic similarity, which is the degree of similarity between the meaning of the AI's answer and the correct answer. The study compared the performance of ChatGPT with human-written responses on four different topics:
· Literature
· History
· Legal & Ethical
· Sports
To ensure consistency in the results, the researchers conducted their test 30 times, measuring two important components: Accuracy and Response Time.
In each of the individual subjects, ChatGPT gave much more detailed and accurate answers than short, human-written baseline answers:
· Legal & Ethical: ChatGPT did the best in this category with an average accuracy of 84.25% (and an accuracy of 90.33% on some test templates). The average baseline answer for humans was approximately 43.83%.
· Sports: The accuracy of ChatGPT was 64.25% on average, while the human baseline was 34.50%.
· Literature: ChatGPT scored an average accuracy of 57.08%, whereas the average human performance was 16.25%.
· History: For the subject of history, ChatGPT achieved an average accuracy rate of 46.92% compared to 15.50% for humans.
While ChatGPT provided much more detailed and accurate explanations, it took much longer to generate them:
· Since the prompts were concise, Human responses took about 16-21 seconds, on average, per prompt.
· The average time for ChatGPT answers was approximately 38 to 48 seconds, as it provided lengthy, informative answers.
Lesson: More processing time is needed for higher accuracy and detailed reasoning. Unlike humans, ChatGPT is more interested in depth than speed.
If you rely on AI for study or work, here are the main practical takeaways from the study:
· Great for Complex & Structured Subjects: ChatGPT shines in rule-based or structured domains like law and ethics, making it a great study aid or research assistant.
· Double-Check Historical Facts: Because open-ended or highly specific historical details can lead to inaccuracies ("hallucinations"), always double-check AI-generated history facts.
· Human Judgment Still Matters: The authors emphasise that while ChatGPT is an amazing helper, AI should support, not completely replace, human critical thinking and expert verification, especially in education, medicine, and legal fields.
For further reading and to explore the original research paper, check out the following resources:
· Original Research Article: Aladdin, A. M., Muhammed, R. K., Sardar Abdulla, H., & Ahmad Rashid, T. (2026). ChatGPT: Precision Answer Comparison and Evaluation Model. InfoTech Spectrum: Iraqi Journal of Data Science, 3(1), 74–90. https://doi.org/10.51173/ijds.v3i1.60
· YouTube Channel: https://www.youtube.com/@tarika.rashid2096