Why developers may still need AI training

AI

By Mike Agoya

Published: 2026-10-06T17:39:19 · Updated: 2026-10-06T15:39:19Z

Why developers may still need AI training

Using an AI coding tool can feel easy enough to learn on your own. You describe what you want, refine the prompt when the answer misses, and keep going until the code looks right. For many developers, that can make formal AI training sound unnecessary.

Andela’s latest study suggests structured training can still help. The company assessed 41 professional engineers in Kenya before and after two weeks of structured training on OpenAI Codex, and average scores rose from about 69% to 81%. Nearly three in five participants improved.

What changed was not whether the engineers understood that AI could write code. Most already did. The bigger gains came from learning how to work with an agent inside a real development environment: what it can access, what context it needs, where instructions should live, what limits apply and how to define whether the work has actually been completed correctly.

That showed up clearly in the assessment. On a question about sandbox and network restrictions, correct responses rose from 51.2% to 90.2%. Repository-guidance hygiene improved from 58.5% to 90.2%, while scores on verification criteria climbed from 61% to 87.8%.

Those are the parts of AI coding that are easy to overlook when the visible interaction still looks like prompting. Once an agent is editing files, running commands and working across a repository, the job becomes less about asking for code and more about giving the system the right environment to work in.

Those numbers get closer to the practical difference between using an AI chatbot while coding and working with an agent that can operate inside a software project. Codex can inspect files, edit code, run commands and work through multi-step tasks, which means the engineer has to do more than describe the desired feature. The agent needs enough of the surrounding system to make sensible decisions, but not so much irrelevant context that its instructions become noisy. It also needs boundaries. If a task depends on an external API that the execution environment cannot reach, rewriting the prompt will not fix the problem.

Andela tested exactly that distinction. One assessment scenario described a developer whose agent repeatedly produced code that matched the prompt but did not follow patterns already established in the repository. Participants had to identify whether the problem was the model, the complexity of the task, the wording of the prompt or missing context from the existing codebase. Even before training, 90.2% chose the correct answer: give the agent examples of the patterns it is expected to follow. After training, that moved only slightly to 92.7%. By comparison, questions that required engineers to understand how the agent's environment actually behaves produced much larger gains.

Some of the harder decisions remained hard

The programme did not produce improvement across every part of the assessment. In three areas, the share of correct answers fell after training, and two of those questions dealt with choices that can look reasonable in several different implementations.

One scenario asked engineers how a team should make an incident-response playbook available to Codex when the document changes frequently. Andela's preferred answer was to move the durable guidance into a version-controlled Markdown file in the repository and reference it from the agent's instructions, keeping the information close to the environment in which Codex is working. Before training, 43.9% selected that answer. Afterwards, only 36.6% did. Another question asked how Codex should be integrated into a continuous-integration pipeline so every pull request receives an automated security review that can block a merge. The expected approach used the Codex SDK from the CI pipeline, but correct responses fell from 46.3% to 41.5%.

Engineers can understand the principle that an agent needs context and still struggle with the mechanism used to provide it. A repository instruction file, an MCP connection, a reusable skill and an SDK integration can all expose information or capabilities to an agent, but they solve different problems. Choosing between them requires knowledge of how the development environment is structured and what needs to remain persistent, automated or available during execution.

The strongest participants had less room to improve. Eleven engineers who entered the programme answering six or fewer questions correctly gained an average of 3.45 questions, moving from 5.36 to 8.82 out of 12. Another 11 who began with seven or eight correct answers gained 1.64 questions on average. Among the 19 participants who started with nine or more correct answers, the average gain was just 0.05 questions; they finished at 10.42 out of 12.

A participant starting at ten or eleven correct answers simply has less room to move than someone starting at five. But it also shows why a single "AI training" programme will not land the same way across an engineering team. Developers who already understand agent workflows may need less introductory material and more practice making architectural decisions around tools, permissions, automation and verification. Someone encountering those workflows for the first time has a much larger set of operational concepts to pick up.

The engineers came through the Andela x OpenAI Codex Accelerator, a programme announced in April and limited to Kenya-based talent. Andela described it as a three-week course requiring roughly six hours a week. The first two weeks combined self-paced learning with four live workshops covering Codex setup, context engineering, debugging, refactoring, planning and agentic workflows. A separate third week was reserved for participants to build and submit capstone projects, with the strongest projects scheduled for demonstration at an Andela and OpenAI meetup in Nairobi.

The research paper examines the instructional part of that programme, not the capstone projects. Participants took the post-training assessment immediately after the two teaching weeks and before receiving access to the capstone, allowing Andela to compare their answers before and after the same body of instruction. The company says the study was conducted as an ancillary evaluation of the learning programme, not as a controlled experiment designed to prove causation.

The study used the same 12-item scenario assessment before and after training, had no control group and measured whether participants could recognise the preferred approach to a set of Codex-related problems. It did not track whether their code contained fewer bugs, whether projects shipped faster or whether agent-generated changes were more secure in production. Andela explicitly separates those claims from what its results can establish.

These engineers already understood much of the theory behind coding with AI before the course began. The biggest gaps appeared when that theory met the development environment: permissions, repository context, verification and the rules governing what an agent can actually do.