Python can beat C++ when the C++ chases pointers.
Here is the catch. How your data sits in memory can matter more than which language you use. Once the data outgrows the processor caches, a program that reads records scattered across memory waits on memory for almost every record. This is a well-known effect, and the page shows its size on one real machine, with code you can rerun.
One set of records in three layouts
This C++ program adds up one field of every record. Add more records and watch the time per record.
Now the catch
We gave five languages one job: walk through 33.5 million items and read each one, doing nothing else. Each bar shows the time to read one item. C++ appears twice, once reading a flat array and once reading separate objects scattered through memory.
A Python loop that reads a flat array is about twice as fast as a C++ loop that must chase pointers. Memory layout matters much more than language. Fix the memory layout, and C++ reads a flat array using 0.9 ns per element, about 24 times faster than Python.
Not a study. The effect is textbook. Ulrich Drepper's paper “What Every Programmer Should Know About Memory” is the standard reference. The page shows how large it is on one machine.
Machine: a rented eight-core virtual machine, which hides its real cache layout. Each loop ran on one pinned core, and each figure is a median of five repetitions from one hand-written loop.
First dial: C++ adds one field across N records of 32 bytes. The three layouts are a contiguous array of structs, separate heap objects visited in random order, and a linked list at random addresses. The stops are 512, 32,768, 1,048,576 and 33,554,432 records. Java and Go follow this pattern: at 1 GB the random-order objects take 60.6 ns (Java) and 18.7 ns (Go) per record, against 3.2 and 3.3 ns for a flat array.
Second block: a read with no arithmetic, at 33,554,432 elements. Python reads array('q') with for x in a, PHP uses foreach over a packed array, Ruby uses a[i] in a while loop, Node reads a Float64Array, and C++ reads either an array or scattered heap objects. The totals include loop cost. An empty loop shaped like the read took 18 to 19 ns in Python and 0.2 ns in C++.
Data: honest-code-traces/measured/layout and measured/layout-raw.