Given the topic complexity and the length of this article I have split it in 3 three different blog-post:
Step 1: What is PGO
PGO (Profile-Guided Optimization) is a two-pass compilation technique.
First, we build the program with instrumentation that adds counters to every branch, loop, and call site.
Next, we run it against a representative training workload so these counters can record which code paths are actually hot versus cold.
We then recompile the program using that data to drive the compiler's decisions.
With this profiling data, the compiler inlines hot call sites more aggressively and lays out hot paths as straight-line, fall-through code while pushing cold paths out of the way.
Finally, it physically reorders functions in the binary so frequently interacting hot code sits close together for better instruction-cache locality, and converts frequently resolved indirect calls into direct ones.
The result isn't "make the training run fast" but rather "reshape the entire binary's layout around real execution patterns," which is why the quality and representativeness of the training workload matters so much to whether PGO actually helps.
It is nice to be wrong
Not so far ago I was wondering if having Percona Server compiled with PGO default is a good idea or not. Then I started to do some tests and I end up with this:
If I compare non PGO with PGO release I can see an optimization, minimal below 5% but is there.
If I compile MySQL using PGO and use as sample the recording of a specific test say sysbench-tpcc my compile will always be slower than my compile without PGO, no mater how many threads I use during recording, I tried from 128 to 1024.
Given my understanding of PGO was that it should be the other way around I was a bit disoriented. So I decided to read a bit and get a better understanding of what PGO really means/does and if it makes sense or not.
