Comparison of error numbers in code written by humans and AI

Researchers from CodeRabbit analyzed 470 pull requests (350 created by AI, 150 written manually) in open projects on GitHub and concluded that changes generated by AI assistants contain 1.7 times more significant defects and 1.4 times more critical issues than manually written code. On average, AI-generated pull requests had 10.83 problems, while the manually created changes had a rate of 6.45.

When examining specific categories of issues, the AI-generated code had 1.75 times more logical errors, 1.64 times more quality and maintainability problems, 1.56 times more security issues, and 1.41 times more performance problems. Additionally, it was noted that the AI-generated code has a 1.88 times higher likelihood of improper password handling, a 1.91 times higher risk of unsafe access to objects, 2.74 times higher chances of cross-site scripting (XSS), and a 1.82 times higher risk of unsafe data deserialization. Meanwhile, the human-written code had 1.76 times more spelling mistakes and 1.32 times more testing-related errors.

Comparison of error numbers in code written by humans and AI
Comparison of error numbers in code written by humans and AI
Comparison of error numbers in code written by humans and AI

Some other studies:

  • In a study conducted in November by Cortex, it was noted that compared to last year, the number of pull requests created by a single developer increased on average by 20% due to the use of AI, but the number of issues in pull requests grew by 23.5%, and the rejection rate for changes increased by about 30%.
  • An August study by the University of Naples concluded that AI-generated code is generally simpler and more uniform but contains more unused constructs and embedded debugging statements, whereas manually written code is structurally more complex and contains more maintainability issues.
  • A July experiment by the METR group showed that AI assistants do not speed up task resolution, but rather slow it down, even though participants subjectively felt that AI had accelerated their work.
  • A study from Monash University in January states that GPT-4 generates more complex code, requiring refinement for ongoing maintenance, but performs better in passing tests.

Source: opennet.ru

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster