Distributed tracing backend
Design the ingestion and storage backend for a distributed-tracing system that collects spans from thousands of microservices producing 10M spans/sec. A single user request can produce a trace of hundreds of spans across services; engineers need to query a full trace by id within seconds, and run analytics ('p99 latency of checkout last hour'). Storing every span is too expensive at this volume. How do you ingest, sample, assemble traces, and store them so both trace-lookup and aggregate queries are fast and affordable?
What a strong answer looks like
Clarify scale and constraints first. Propose a clean component breakdown, then go deep on the hard parts (data model, bottlenecks, consistency, failure modes) and name the trade-offs you are making.
Clarify5:30 left
Estimate5:30 planned
Design16:30 planned
Deep dive13:30 planned
Failure9:00 planned
Which questions mattered is sealed until you submit. Telling you now would just be handing over the edge cases.
Run or narrate your approach, then ask the coach.