100,000+ Verified Real-World Prompts

Company Question Bank

Browse real questions candidates were asked at leading technology companies, ranked by frequency with verified model solutions and pitfall warnings.

GoogleSenior Software Engineer•system design•L5 / Senior

Design a globally distributed rate limiter that handles 10M requests/second with sub-5ms latency and prevents cascading failure.

98.4%
Frequency

High-Impact Answer Approach

Token Bucket with localized Redis in-memory caches, distributed consensus via local batching, and fallback degraded mode.

Key Talking Points for Teleprompter

  • Token Bucket vs Sliding Window Log: Token Bucket provides the optimal memory footprint and handles bursty traffic cleanly.
  • Local in-process memory cache synchronization using Redis cluster with asynchronous batch increments.
  • Degraded fail-open architecture: If rate limiter cluster experiences network partition, allow traffic through with metrics alert.
Reference Solution (Python 3):Optimal Complexity
class TokenBucketRateLimiter:
    def __init__(self, capacity: int, refill_rate_per_sec: float):
        self.capacity = capacity
        self.refill_rate = refill_rate_per_sec
        self.tokens = capacity
        self.last_refill = time.time()

    def allow_request(self) -> bool:
        now = time.time()
        delta = now - self.last_refill
        self.tokens = min(self.capacity, self.tokens + delta * self.refill_rate)
        self.last_refill = now
        if self.tokens >= 1:
            self.tokens -= 1
            return True
        return False
Critical Pitfall to Avoid: Avoid doing a remote network roundtrip to a single centralized database on every HTTP request.
MetaFull Stack / Systems•algorithms•E5 / Senior

Implement an LRU (Least Recently Used) Cache with O(1) time complexity for both get and put operations.

96.8%
Frequency
AmazonSoftware Development Engineer•behavioral•SDE II / III

Tell me about a time you took a calculated risk and made a two-way door decision that didn't go as planned.

94.2%
Frequency
OpenAIAI Systems Engineer•system design•Senior

Design a real-time streaming inference gateway for large language models that minimizes Time-To-First-Token (TTFT).

97.1%
Frequency
AppleCore OS / Embedded•algorithms•ICT4

Find the median of two sorted arrays of sizes m and n in O(log(min(m, n))) runtime complexity.

91.5%
Frequency
MicrosoftPrincipal Architect•system design•L65 / Principal

Architect a globally resilient multi-region database replication scheme under Azure SQL / MS SQL Server.

93.7%
Frequency