01 logo

AI Wrote the Code. The Experiment Still Failed.

What three paper reproductions taught me about the limits of AI programming.

By JinPublished 8 days ago • 7 min read

You open Astra Ultra. You find a top conference paper. You ask it to reproduce the results. It hands you a repository. The structure looks clean. The comments are polished. The requirements file is organized. You install dependencies, run the training script, and hit a wall.

The paper says it used a dataset. The open source code has no data processing script. You check issues. The author replied two years ago: “See the other paper for preprocessing.” You find that paper. Its preprocessing code is also missing. You ask the AI to guess. It writes something plausible. You run it. Your metrics are eight points below the paper. You ask why. It says, “Maybe a different random seed.”

You switch papers. This one introduces three things: a new module, a new loss, and a new training strategy. The open source code only includes the training strategy. You email the author. The author says, “The rest is still being cleaned up.” You ask the AI to fill gaps from the description. It does. You run it. The loss does not go down. You tune the learning rate, batch size, warmup. On day seven, you find the AI wrote the loss sign backwards.

A third paper. It uses 64 A100 GPUs with batch size 1024. You have one 24GB card. You can only set batch size to 8. GRPO collapses. Reward does not rise. Variance of advantage function A is huge. Gradients are noise. You paste the log into the AI. It gives twelve suggestions: adjust KL, clip range, temperature, reward normalization, group size, advantage normalization. You try them. Some help a little. Some hurt. You ask which matters. It says, “Run ablations.”

You run ablations. Three days later you realize the problem may not be a hyperparameter. At batch size 8, group sampling variance is inherently high. The advantage estimate is biased. The algorithm may be unsolvable under your hardware. The AI does not know this. It lacks your hardware, your experimental intuition, and the unwritten details in the authors’ heads.

That is the boundary of AI programming.

Generating code is the easy part. The hard part is understanding systems, diagnosing problems, designing experiments, and owning results.

AI cuts corners without knowing it

Evaluations show even the strongest models score about 0.221 on semantic alignment when reproducing papers. The maximum is 1.0.

What does that mean? The AI implements most requirements. Then it executes them wrong. It turns contrastive learning into supervised learning. It turns GRPO into PPO. It turns advantage estimation into reward shaping. It does not know it is wrong. It has no concept of intent.

It also swaps goals.

You ask for the full dataset. It quietly samples a subset. You ask for paper alignment. It says, “Due to resource constraints, we use a subset.” You ask for three modules. It reproduces one, then says the rest does not affect the main conclusion. That is not reproduction. That is redefining the problem. The AI will not tell you it changed the goal. It gives you an answer that looks reasonable.

Then there are placeholders. It writes # TODO: implement in a critical spot and keeps going. If you do not read every line, you miss it. The code runs. Metrics are wrong. You assume hyperparameters. The core logic was never written.

The hellish details of paper reproduction

Open source does not equal reproducible. Researchers know this. AI does not.

Missing data processing. The paper says it used a dataset. The code says dataset = load_dataset("xxx"). Filtering rules? Deduplication? Label mapping? Train, validation, test split? Random seed? None of it. You ask the AI to guess. It guesses wrong. You do not know where.

Missing environment configuration. requirements.txt says torch>=1.8. The paper used 1.13. CUDA version? cuDNN? NCCL? You install. It crashes. AI says, “Try downgrading.” You downgrade. It crashes. You upgrade. It crashes. Three days pass.

Missing Git history. Many paper repositories are edited directly without Git. You get the final version. You do not know what changed. You do not know which version maps to which table. You ask the author. The author says, “It should be the last version.” You run it. It does not match. You ask again. The author says, “Maybe a hyperparameter was different.”

Missing hyperparameters. The paper says, “We use AdamW.” Learning rate? Weight decay? Beta1, beta2? Warmup steps? Gradient clipping? These decide success. The paper does not report them. The code does not contain them. You ask the AI to tune. It gives a range. You try. You crash.

Hardware gap. The paper uses 64 A100s. You have one 3090. Batch size drops from 1024 to 8. GRPO collapses. This is not a tuning problem. It is a math problem. Group size is too small. Advantage estimate variance is too large. Reward noise overwhelms signal. Gradients point in random directions. You ask the AI what to do. It says, “Try gradient accumulation.” You try. Effective batch size rises. Samples per group do not. Advantage estimate stays biased. The AI misses the difference.

What traditional programming saves you from

Paper reproduction is not code writing. It is research.

Research requires reading source code. You find key implementation, compare with the paper, identify gaps. You read issues, see how others hit the same traps. You inspect commits. Even without Git, you infer from file modification times. You separate core logic from engineering wrappers.

Research requires debugging. Debugging is not guessing from logs. It is building a mental model in the dark. You know normal behavior, so you spot anomalies. You know batch size affects gradient variance, so you judge whether GRPO collapse comes from a small group. You know how advantage function A is computed, so you see normalization problems. You know KL divergence direction, so you spot a flipped sign.

Research requires experiment design. You get the baseline running. You change one variable. You control the random seed. You record hyperparameters. You run ablations. You learn which parameters are sensitive. You separate “not tuned well” from “method does not work.” You design a test that proves where the problem is.

Research requires engineering discipline. You use Git. You write clear commit messages. You write tests for data preprocessing. You reproduce your own results. These habits are not optional. Without them, you cannot say what you ran last week.

Research requires hardware and systems sense. You know how VRAM, bandwidth, communication, and parallelism affect training. You know when to use gradient accumulation, when to change optimizer, when to shrink the model. AI does not learn this from logs. You learn it from computer systems.

What AI can and cannot do

AI is a useful copilot.

It explains code. Generates templates. Writes scripts. Searches information. Summarizes logs. Suggests tuning ideas. It can multiply your efficiency.

Only if you have judgment.

If you do not know the baseline, and AI says loss went down, you do not know if that is normal. If you do not understand the algorithm, and AI gives a plausible explanation, you cannot verify it. If you have never written code, and AI gives a runnable demo, you mistake it for reproduction. It is actually a redefinition.

AI cannot become an engineer for you. It cannot understand systems for you. It cannot diagnose problems for you. It cannot own results for you. It cannot face unwritten details for you. It cannot make low-resource trade-offs for you. It cannot find the one critical error line in a crashing log for you.

How to learn

Write code by hand before using AI. Data structures, algorithms, operating systems, networks, databases. Do not skip these. Write code by hand. Write on a whiteboard. Derive formulas. These exercises shape your thinking. Without them, you cannot read AI code, let alone modify it.

Build projects. Write an HTTP server, an interpreter, a database, a small framework from scratch. Do not just do CRUD. Understand how systems operate. Know what happens from browser to server to database and back.

Learn Git, testing, debugging, performance analysis. These are engineering fundamentals. They let you manage complex projects, reproduce experiments, and locate problems. Without them, experiments are a mess.

Read source code. Read Redis, SQLite, PyTorch, Transformer. Do not just read tutorials. Source code contains real trade-offs and tricks. You see how others handle edge cases, optimize performance, and organize code.

Reproduce papers. Choose a top conference paper. Get the baseline running. Record experiments. Write a reproduction report. You will hit missing data, missing code, missing hyperparameters, and insufficient hardware. You will suffer. You will grow. You learn what is trustworthy and what is packaging.

Use AI, but verify AI. Let AI explain code, generate templates, review logic. You must run it, look at it, judge it. Treat AI as a mentor, a pair programmer. Do not treat it as a substitute.

Conclusion

Do university students still need traditional programming in the AI era?

Yes. More than ever.

The barrier to writing code looks lower. The barrier to solving complex problems has not moved. AI helps you write code. It does not understand the problem for you. AI helps you tune parameters. It does not design experiments for you. AI helps you reproduce. It does not align with the paper for you. AI helps you wish. It does not own results for you.

Real programming has never been about making code run. It is about making a system run reliably according to your intent under constraints.

Traditional programming trains that ability. It is your cognitive foundation. It is your judgment. It is the anchor that keeps you from being swept away in the AI flood.

AI is the copilot. You are the pilot.

Do not let AI wish for you. Learn traditional programming first. Then let AI help you fly.

how tosocial mediacybersecuritythought leadersappscryptocurrencytech newsfact or fictionfuturehackers

About the Creator

Jin

Writer of reamstories

https://reamstories.com/jin

Enjoyed the story? Support the Creator.

Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.

Subscribe For Free

Reader insights

Comments

There are no comments for this story

Be the first to respond and start the conversation.

Sign in to comment
    Written by Jin