There are some good examples of why that's not sufficient in Russ' blog post, which is worth reading. They include performance regressions that may not be immediately visible, understanding cache utilization, and better optimizing the toolchain for real-world machine configurations. It's worth reading his article for a better understanding of the goals here:
https://research.swtch.com/telemetry-intro
(Disclosure: I work on developer tools at Google.)