He Didn’t Say “I’m Good at Databases.” He Said, “I Blew One Up Before.”
Alex didn’t talk about system design. He sketched a pothole, a lock he crossed out, and the 2 a.m. cron job that taught him to ask about the business first.

The whiteboard hadn’t been wiped clean. Down in the left corner, half an ER diagram from the last candidate was still there, and in the upper right corner, a DingTalk reminder about a food delivery clung to the board. Nobody had peeled it off. I didn’t erase the half‑finished drawing. I just handed him the marker.
“Assume we’re building an order system. It needs to support our current daily volume, doubled. You draw first.”
His name was Alex. Six years of experience on his résumé, three of them on an e‑commerce middle platform. He didn’t start drawing right away. He laid the marker down in the whiteboard tray and asked a question first: “What’s the current volume? And when does the doubling happen—during a peak‑sale spike, or as a steady daily average?”
I wrote the numbers in a corner: 120,000 orders a day, doubled means 240,000. Peak volume is three times the daily average.
He nodded, picked up the marker, and wrote at the very top of the board: Peak QPS estimated 4,200. Then he drew a long horizontal line that stretched almost the entire width of the board.
“This layer. I’m putting only the gateway and rate limiting here. No business logic.”
Beneath the line he drew a box and wrote Order Service. From that box he extended two arrows: a solid line pointing to Inventory Service – synchronous deduction, and a dashed line pointing to Message Queue – asynchronous notification.
He stared at those two lines, his thumb popping the marker cap off and clicking it back on (off, on, off, on) four or five times. Then he reached over and crossed out the solid line.
“If inventory deduction is synchronously coupled here, a spike in order timeouts will drag the inventory service down with it.” He pointed at the peak number I’d written. “At 4,200 it won’t die. At 8,000 it definitely will. I’ve already filled in that pothole before.”
He didn’t say, “I have a lot of experience.” He said, “I’ve already filled in that pothole before.”
I followed up: “Then with async, how do you make sure the inventory the user sees is accurate?”
He turned the marker sideways and drew a small box in the margin of the board, writing two characters inside: Cache. Then right next to it he drew a short dash and wrote 5 min.
“Expires in five minutes. The inaccuracy the user sees won’t last longer than five minutes. If you want stronger consistency, add a Redis distributed lock—” He paused, drew a tiny lock symbol above that small box, and then crossed it out himself, “but I wouldn’t recommend it. The maintenance cost of a lock far outweighs the experience loss from those five minutes.”
He set the marker down and looked at me. “It’s a trade‑off.”
That word, trade‑off—he said it on his own. I didn’t prompt him. Even before saying it, he had drawn an alternative himself and crossed it out.
I picked up my water bottle, took a sip, and asked, “Then for your gateway rate limiting, how do you set the thresholds? Is that 4,200 you wrote the system’s absolute upper limit, or do you want to clamp it somewhere lower?”
He picked up the marker again and wrote two numbers next to that long horizontal gateway line: 3,800 and 4,500.
“3,800 is the soft limit. Beyond that, only VIP users and internal whitelist requests get through. 4,500 is the hard limit. Anything more gets rejected outright and sees a queue‑it page.” He tapped the tip of the marker on the 3,800. “I didn’t pull this number out of the air. I assumed the downstream order service has a connection pool of 200 connections, average response 50 milliseconds, worked out the throughput ceiling, and then applied a twenty‑percent safety margin. When you haven’t done enough load testing, twenty percent off is the safest bet.”
He didn’t say, “I’ve done load testing.” He said, “When you haven’t done enough load testing, twenty percent off is the safest bet.”
From the whiteboard tray he picked up another marker—a blue one, nearly out of ink—and next to the order service box he drew another small box, writing inside it three characters: State Machine. Then he began drawing arrows: Pending Payment, Paid, Picking, Shipped, Completed, Canceled, Refunding. Between every two states he connected a thin line and wrote a tiny trigger condition along it. His handwriting was already small. When he got to Picking → Shipped, the blue marker ran completely dry. He used his fingernail to score a groove along the line and said, “This transition has to wait for a callback from the WMS. You can’t just poll with a scheduled task.”
Then, below the state machine diagram, he wrote a line: State transition tables: one table for current state, one table for history. Current state table only allows primary‑key queries. No range scans.
He turned to me. “History table, you can range‑scan it. But if the current state table ever gets range‑scanned, during peak traffic the disk IO will redline instantly. I blew one up before at my last company when I mixed presale orders and spot‑sale orders into the same scan.”
He didn’t say, “I’m good at database optimization.” He said, “I blew one up before when I mixed presale and spot‑sale orders.”
I asked him to circle the Refunding state on the board. “What do you think matters most with this state?”
He tapped the Refunding label with the blue marker, thought for a moment, and drew a small funnel beside the state, mouth pointing downward. “A refund isn’t just reversing a payment. If the user requests a refund before shipment, you need to unlock the inventory, release the coupon, and return the loyalty points all at the same time. Those three actions have to be inside the same transaction, but the coupon system and the points system are separate services.”
From the funnel he drew three arrows reaching out, labeling them Inventory, Coupon, Points. At the point where the three arrows converged, he drew a large curly brace and wrote TCC.
“Compensating transaction. Try first, then Confirm. If Confirm fails, retry. After three retries, if it still fails, drop it into a manual work queue. You cannot let a machine get stuck dead in a refund loop. Getting stuck is worse than losing money.”
He set the marker down. His palm had picked up blue smudges; he wiped it twice on his pants, not seeming to care. He kept pointing at the funnel. “The most fragile link here isn’t the inventory, and it’s not the points. It’s the coupon. Coupons often have campaign time windows. If the refund window lands right after a campaign ends, the coupon might already be invalid. At that point, what do you return to the user? The coupon itself, or cash equivalent? This isn’t a problem you can solve with a technical choice. You have to confirm it with the product team right at the start.”
He wrote the words Confirm with product team in the bottom‑right corner of the whiteboard—a spot no one would ever deliberately look at. But he wrote them with force, and on the very last stroke the blue marker broke off a small chunk of its tip.
During the second round of interviews, he told me about a project. It didn’t sound like a project. It sounded like a disaster site with the lights still on at two in the morning.
The original system would freeze up every few days, always around 2 a.m. The monitoring curve would spike from a normal 60 milliseconds to 11 seconds in an instant. Then the whole cluster would start avalanching. Even login wouldn’t work. Nobody could reproduce it, and nobody could isolate it. On his third day, he copied every single cron expression for every scheduled task in the system onto a sheet of A4 paper—not printed, handwritten in pencil, all seventeen of them. He slapped that sheet onto the desk and went down the timeline, one row at a time.
“I drew a timeline and marked every cron trigger moment on it. I found one statistics task—scheduled precisely at 2 a.m. every day—and it was overlapping almost exactly with the database’s master‑slave switchover window.”
He picked up a sticky note from my desk, drew a line, wrote 2:00 stats task at one end, 2:00 master‑slave switchover at the other, connected the two points, and drew a red arrow, writing Deadlock at the arrow’s tail.
“At the moment of switchover, the master database becomes read‑only. That big stats SQL hit the slave database, and the slave couldn’t handle it at all. Because it was triggered by cron, nobody knew that particular SQL was running. The logs were full of timeouts. The SQL itself wasn’t there.”
He crossed out the 2:00 on that sticky note and wrote a new time beside it: 3:15.
“I changed the cron to 3:15, dodging the switchover window.” He paused, and then added a line at the bottom of the note. “And then I added a monitoring alert: if slave replication lag exceeds 5 seconds for two consecutive checks, send an SMS immediately.”
He pushed the sticky note toward me. The bottom‑right corner of the paper had a crease from the pressure of his pen. He said, “Finished this whole pack that week.” He pulled a box of stomach medicine from his pocket. The foil blister pack was nearly empty; only one last pill hadn’t been popped out. He shook it. The pill made a tiny sound inside the foil.
The whole room laughed. He didn’t say, “I’m good at troubleshooting.” He said, “Finished this whole pack that week.”
I asked him, “After you added that alert, did it ever actually go off?”
He nodded. “Twice. The first time was three months later. Another team deployed a new service and also set their cron for 2 a.m. The second time was on Double Eleven.”
“So what did you do?”
“The first time, I pulled that team over to our desks, cracked open a Coke, and told them, ‘Hey, 2 a.m. to 4 a.m.—that window is mine. Could you guys switch?’ They moved theirs to 5 a.m.”
“And the second time?”
“The second time I didn’t change the cron. On the early morning of Double Eleven, I temporarily lowered the alert threshold from 5 seconds down to 2 seconds, and escalated the notification from SMS to a phone call. Then I set an alarm for 1:50 a.m., sat down in front of my computer ten minutes early, and just stared at the monitoring curves. I didn’t stop staring until 2:15.”
He pulled out his phone, scrolled to his alarm list, and showed me. There was a daily repeating alarm. The label didn’t say “Wake up.” It said “Switchover window.” “Later on, the master‑slave switchover architecture got optimized, but I never deleted this alarm. I kept it.”
The screen protector on his phone had a small crack in the lower left corner. He hadn’t replaced that either.
For the last ten minutes, I wiped the whiteboard clean—this time I did it myself—and then asked him, “If you had to build this order system all over again, what would you do differently during the design phase?”
He thought for a long time. His hand spun the marker unconsciously, four or five full rotations. The cap fell on the floor. He bent down, picked it up, spun it another half turn, and then finally spoke. “I would first ask the business team why that statistics task had to run in the middle of the night.”
“Why?”
“Because the business team didn’t actually care about the exact time. They only cared about seeing the numbers before they got to work. I later moved the task to 5 a.m., and it perfectly avoided the switchover. Why didn’t I ask back then? Because I walked in assuming the requirement couldn’t be changed.” He placed the marker back in the tray and wiped his nose lightly with the back of his hand. “I’ve got a bad habit called ‘diving into technical details too early.’”
He didn’t use a narrator’s voice to declare, “I have the ability to reflect.” He just gave his bad habit a name and wiped his nose with the back of his hand.
The interview ended. He put the stomach medicine back into his pocket. He also folded the sticky note, ready to toss it into the trash. I reached out and took it from him. “I’ll keep this one,” I said. He paused for a second, then smiled.
As he stood up, he knocked over the bottle of mineral water on the table. Some spilled out, forming a small puddle that dripped over the edge of the desk. He righted the bottle, screwed the cap tight, and used his own sleeve to wipe the table. He didn’t just wipe once and stop. He kept wiping until every trace of water was gone. The cuff of his sleeve was soaked. He didn’t seem to care at all.
I closed my notebook and said to him, “Contact HR directly. I’ll write the email in a moment.”
He walked to the door and then looked back at the whiteboard. It was clean now. The half ER diagram from the previous interview in the lower left corner was gone, too. What he looked at was the bottom right corner—the spot where he’d written Confirm with product team. The writing was still there, a faint shade of blue, and the tiny gap from the broken marker tip was still there.
When he pushed the door open and stepped out, the motion‑sensor light in the hallway came on. But the direction he walked, the light didn’t follow. His silhouette stayed in the light for two seconds and then melted into the dark.
I unscrewed the bottle he had tightened, took a sip of water, and opened my email.
About the Creator
Jin
Writer of reamstories
https://reamstories.com/jin
Enjoyed the story? Support the Creator.
Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.
Comments
There are no comments for this story
Be the first to respond and start the conversation.