Why Every Realistic Video Game Still Feels Fake
It’s not 8K. It’s not ray tracing. The human brain ignores bad shadows but catches a fake blink in a heartbeat, and that’s the detail game studios keep missing.

The More Realistic, the More Fake: The Perceptual Paradox of Realism in Games
The human eye is almost blind to a wrongly drawn shadow, yet it can spot a fake blink in an instant.
In 2005, Ostrovsky, Cavanagh, and Sinha ran experiments on how people perceive illumination inconsistencies in scenes. Their conclusion used the phrase “remarkably insensitive.” People are insensitive to lighting errors. Nightingale et al. narrowed the question in 2019 to errors in shadows and reflections. The overall detection rate was poor. When the error was subtle, participants were close to guessing.
The same eye, looking at Sam’s face in Death Stranding 2, can tell it is not a real person. After the UK’s Online Safety Act came into force in 2025, some players used the photo-mode Sam to pass facial age verification. It worked. PC Gamer reproduced it with the otter-hat version. It worked too. The machine could not tell. The human could. The difference appears once the face moves: blink frequency, microsaccades, gaze focus, whether lip shapes match phonemes.
Put those two facts together, and you have the answer to why game graphics keep getting more realistic while players keep getting better at spotting the fake.
Resolution ran out of road years ago
The resolution limit of the human eye is often given in engineering as 60 pixels per degree at the center of gaze. A 2024 study on the resolution limit of the eye drew that line more finely with measured data: at 20 degrees eccentricity, the median observer has only about 22 pixels per degree, and the 95th percentile is about 35 pixels per degree.
Convert that to a living room. Watching a 55-inch TV from 3 meters away, most people can no longer distinguish individual pixels at 1080p. Going from 4K to 8K yields almost nothing.
The GTA VI extended demo is a useful counterexample. Digital Foundry called it “perhaps the most impressive showcase we’ve ever seen in real-time rendering.” It ran on a standard PS5 at native 1440p and 30 fps, with no FSR and no upscaling. The image that stunned the room was sitting at a resolution from seven years earlier. What carried it was ray-traced global illumination and reflections.
Pixels are already surplus. TV makers still sell 8K. Game UI still offers native 4K. But the bandwidth on the human side was filled ten years ago.
Lighting: people cannot see when it is wrong
Every pitch about graphical progress talks about ray tracing: Lumen, path tracing, neural rendering. That is where the money goes.
People are insensitive to lighting errors. A shadow points the wrong way. A reflection misses half a face. Most players walk past it.
The ray-tracing money is not wasted. Lighting contributes to realism through another route: it keeps materials alive.
Skin is the clearest example. GPU Gems 3 gives the number: about 94% of light reflected from skin comes from subsurface scattering. Light enters the skin a few millimeters and scatters back out. Only about 6% reflects directly from the surface. Without subsurface scattering, no amount of polygons saves you from plastic. In 2015, Jimenez’s screen-space method brought it down to 0.5 ms per frame at full HD, which is why today’s AAA faces all have it.
UE5’s Lumen solves the same class of problem: indirect light. What color is the corner of a room that no lamp directly hits? Is there light bouncing off the floor under the table? In the past, artists baked this by hand. Lumen makes it fully dynamic. But it has its own breakage budget. In software ray-tracing mode, Lumen Scene covers only about 200 meters around the camera by default. Move the camera fast, and indirect light lags half a beat behind. Nanite does not support skeletal-animated characters or translucency. The scene can have cinematic geometric precision. The person walking through it is still using the old method.
Put those two halves together: lighting holds up the credibility of materials. The thing that makes someone decide “this is a game” within half a second has to be found where the human eye has dedicated hardware.
What gives the game away is faces and motion
In 1973, Gunnar Johansson ran an experiment that has been cited thousands of times: the point-light walker. He attached a dozen or so lights to a person’s joints, turned off the lights, and showed only the moving points to participants. They immediately recognized a person walking. They could judge gender, emotion, and even recognize an acquaintance. A dozen points. No outline, no material.
Later brain-imaging work located this in specific regions. Grossman and Blake’s 2002 experiment in Neuron showed that the superior temporal sulcus activates specifically when viewing point-light walkers, and even the fusiform face area lights up. The brain has dedicated circuits for biological motion and faces, with speed and sensitivity far beyond general vision.
Those two circuits are where games get caught most easily. They never look at polygon counts. They only look at whether motion is right and whether a face is alive.
One concrete piece of evidence on the face side is Death Stranding 2. After the UK’s Online Safety Act launched in 2025, players used Sam’s face from photo mode to pass facial age verification. It passed. PC Gamer reproduced it with the otter-hat version. It passed. The machine could no longer tell whether the face was real. The human still could. The difference appears once it moves: blink frequency, eye microsaccades, gaze focus, whether lip shapes match phonemes. The face-specific circuit calibrates against real people every day.
Cyberpunk 2077 uses JALI, which generates lip sync and facial animation procedurally across ten languages. It is already top of the industry. It also shows why gaze matters: when a character meets your eyes, players post about how the character seems to be talking to them. When Johnny looks down or puts his focus elsewhere, people immediately read “he is not looking at me.” That may not be an animation error, but the brain is far stricter about gaze direction than about shadow direction.
On the motion side, the sensitive point is more mundane: feet.
If animation blending does not lock the contact point of the foot, the character’s foot slides across the ground. The industry calls it foot sliding. The fix is foot IK, pinning the foot to the ground. The problem is that IK has a budget. MoCap Online’s guide gives the numbers: two-bone IK can run dozens of instances at almost no cost; full-body IK on consoles is enough for only one to four foreground characters. So the protagonist walks steadily. The fifth passerby on the street starts ice-skating. The player may not know the term IK, but their superior temporal sulcus sees it perfectly.
Naughty Dog talked at GDC 2021 about The Last of Us Part II’s Motion Matching. The idea is to abandon the animation state machine and search the motion-capture database in real time for the next frame closest to the current pose. It removes the jerk of switching states between running and stopping.
Get motion right and face right, and a rough image can still pass. Get motion and face wrong, and the more detailed the image, the worse it becomes.
The more it resembles a person, the more fake it feels: perceptual inconsistency
Everyone knows the uncanny valley. Masahiro Mori wrote it as an essay in Energy in 1970. It was not formally translated into English until 2012. The original had only a hand-drawn curve: the horizontal axis was degree of human likeness, the vertical axis was affinity, and it dropped sharply just before true human likeness.
For forty years, the curve remained a hypothesis. The actual empirical work has come in the last decade, and the findings are more precise than the original.
MacDorman and Chattopadhyay’s 2016 paper in Journal of Vision, “Familiar faces rendered strange,” used 365 participants. The variables were cleanly controlled. They adjusted the realism of individual facial features: skin, eyes, eyebrows, mouth. Then they compared two conditions: lowering realism across all features together, versus lowering it in only some. Only the latter triggered eeriness: faces with inconsistent realism between features.
The same group reran it with 548 participants and further confirmed that the feeling comes from inconsistency itself, not from whether the face is hard to categorize. Kätsyri et al.’s 2015 review in Frontiers in Psychology went through dozens of uncanny-valley studies. The categorization-difficulty hypothesis found almost no support. The perceptual-mismatch hypothesis did.
That explains why old games did not have an uncanny valley. In the PS2 era, skin, eyes, and motion were all equally fake. The features were consistent. The brain filed them under “drawing” and stopped keeping score.
Today’s characters have pore-level skin and full subsurface scattering, while the eyes still blink mechanically once every thirty seconds. The brain files the skin under “real person,” then immediately crashes into the eyes.
Cross-modal mismatch works the same way. One set of experiments paired a real human face with synthetic speech, and a robot face with human speech. Both combinations were eerier than the consistent versions. Realism is an AND relationship: skin, eyes, motion, sound, physics. If any one falls behind, the rest are wasted.
Predictive coding: realism is a curse
As for why the brain does this, predictive coding gives the simplest account.
Rao and Ballard’s 1999 model says the visual cortex constantly predicts what it will see next from the top down. Sensory input arrives and is subtracted from the prediction. The residual is prediction error, and only that error travels upward. In 2025, Rideaux et al. showed in Journal of Vision that violated expectations are not only noticed but prioritized, with higher representational precision.
Apply that to games and the logic is clear.
The closer a character gets to a real person, the more the brain uses a real-person model to predict them. The real-person model is extremely precise. Blink intervals, weight shifts, eyebrow position during speech. All of it is in there. The more precise the model, the narrower the tolerance. The same small error produces no prediction error on a cartoon character. On a photoreal character, it is a spike.
Realism is a curse. It pushes the player’s tolerance to the floor.
The shape of the uncanny-valley curve has always been disputed. Bartneck et al. compared real photographs with various robots in 2007 and found that real photos were less likable than some toy-like robots. They proposed the “uncanny cliff” model, arguing that the right end of the curve does not rise again. That is even more direct for games: chasing photorealism has no clear finish line. Consistency is what you can actually get.
What 16 milliseconds cuts away
Some of the fakeness in games comes from the budget layer, not from perception.
Pixar’s The Science Behind Pixar project published rendering data for Luca: one scene took 50 hours per frame. A game engine at 60 fps has a budget of 16.6 ms per frame. At 30 fps, it is 33 ms. Fifty hours is 180,000 seconds. Divide by 0.0166, and the two sides differ by roughly a factor of 10 million.
That number explains a lot of the breakage.
Screen-space reflections disappear when an object leaves the frame. Ambient occlusion leaves a dark rim around object edges. Distant vegetation suddenly swaps models. These are things actively cut inside 16 ms. Engineering calls it approximation. Perception calls it a glitch.
DLSS 4 and FSR 4 are the new solutions for this round in 2025. NVIDIA says its Transformer upscaling model is 40% faster than the old CNN and saves 30% VRAM. AMD’s FSR 4 also turns to a machine-learning temporal model for the first time. But frame generation has a concrete cost: fast-moving elements and UI text can smear and flicker, worse when the base frame rate is low.
At the end of that road is Genie 3. DeepMind released it in August 2025. It can generate an interactive world from a sentence of text. The official numbers: it maintains consistency for a few minutes at 720p, 24 fps, with about one minute of memory. Walk back after that, and the scene changes. TechTalks’ review recorded physics errors such as characters walking backward. In the same official demo, a complete road and guardrail degrade into a mud surface within seconds while moving forward, then into open water.
What is lost there is object identity and spatial relations. Resolution is not the issue. The image is pixel-level real. But the underlying system lacks the rigid consistency constraints of a traditional physics engine. The fake is visible faster than with any traditional renderer.
When the world does not respond to you
If we stopped here, the conclusion would be: realism depends on faces and motion.
But one kind of game does not fit that conclusion. Red Dead Redemption 2 came out in 2018. Its resolution and materials are not top-tier today. Eight years later, it is still the world players call the most real. It relies on how NPCs respond to the player. If you are covered in blood, passersby avoid you. If you crouch-walk, the shopkeeper watches you. If an animal is shot in different places, it limps or struggles.
None of that has anything to do with the rendering pipeline. It is the causality of the world.
Chinese-language game criticism has a name for this: the uncanny valley of game mechanics. The more realistic the graphics, the more mechanical the mechanics stand out. AI can be fooled by the same motion over and over. A door can be drawn perfectly and still not open. This is the same model as perceptual inconsistency, only the object of inconsistency has shifted from facial features to the rules of the world.
The image says this is real. The rules say this is fake. The brain is still doing subtraction.
From that angle, stylization is not a fallback. Nintendo’s work, Celeste, and Valve’s decision to move Team Fortress 2 from realistic to illustrative also solved another problem. Valve’s official reason was that the realistic version was hard to read in combat. A stylized image files itself under “drawing” from the start. The brain uses a low-precision predictive model with wide tolerance. A blink half a beat late does not break anything.
Tears of the Kingdom is a generation behind on hardware. It relies on physics rules that are self-consistent from beginning to end. Put a wooden board across two stones, and it becomes a bridge.
Credible and realistic are two different things. Credible comes from internal consistency. Realistic comes from physical accuracy. Games can only get the former.
High pixel metrics, still fake
In medical ultrasound image segmentation, there is a metric called Dice. It measures how much an algorithm’s predicted lesion overlaps with a doctor’s annotation. Large lesions get high Dice scores. Small lesions get mediocre or poor scores. Averaged out, the number looks good, but what gets missed is the small thing that matters most clinically. In 2023, USE-Evaluator in Medical Image Analysis pointed out directly that Dice cannot reflect clinical impact. Pixel-level overlap and a doctor’s judgment are not the same thing.
Image quality assessment quantifies the same problem even more thoroughly. PSNR and SSIM are the two most common pixel-level metrics. A 2023 NeurIPS paper tested their correlation with human subjective ratings. In most cases, Kendall’s τ was only around 0.5, a moderate correlation. The classic counterexample is a slightly blurred image. L1, PSNR, and SSIM all judge it better. The human eye sees it as blurrier. In 2018, Zhang et al. proposed LPIPS, replacing pixel distance with deep-network feature distance, because pixel-level metrics had become so disconnected from perception that the ruler had to be changed.
Resolution, polygon count, and texture precision are PSNR and SSIM. The fake a player spots in an instant is LPIPS. It is perceptual inconsistency. It is one dimension falling behind.
Four variables
Hold everything else fixed and move only resolution: the returns peaked ten years ago.
Hold everything else fixed and move only lighting: materials come alive, but people are insensitive to lighting errors, so the breakage saved here is limited.
Hold everything else fixed and move only motion and faces: every improvement is seen directly by dedicated circuitry. Tolerance is narrowest. Returns are largest.
Hold everything else fixed and move only how the world responds: this is the only variable that can keep a 2018 game the most real one in 2026.
The order of publishers’ budgets and the order of these four marginal returns are exactly reversed.
The human eye is almost blind to a wrongly drawn shadow, yet it can spot a fake blink in an instant. That sentence sounds like a defect in human vision. After looking into it, it is the division of labor.
The brain decided long ago what is worth looking at. The people making games and the people buying graphics cards spent the money where it does not look.
An 8K texture pasted onto a door that will not open. Next to the doorframe is a crack that cannot be entered.
About the Creator
Jin
Writer of reamstories
https://reamstories.com/jin
Enjoyed the story? Support the Creator.
Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.
Comments
There are no comments for this story
Be the first to respond and start the conversation.