Isaac Newton tells us that acceleration is proportional to the accelerating force divided by the mass of the object being accelerated.
This is one part of the equation, but there are 2 other significant parts.
The first is gradient: compared with level track a train will accelerate more slowly going uphill and more rapidly going downhill. Conversely, a train will decelerate more rapidly going uphill and more slowly going downhill.
The other is resistance to motion. The biggest part of this is front end resistance hence this:
Possibly true of MUs if set up to deliver constant performance levels but without that a longer rake will be faster as there is only ever one front end resistance. 12 cars of mk1 EMU was always faster than 1x4 of the same type for example.
But, for hauled trains, rolling stock had different resistances. Part of that was the different aerodynamics of MarkI, MarkII and MarkIII coaches (mixed rakes were in the too difficult pile) and different rolling resistances of different bogies.
Resistances to motion could also be affected by wind direction and speed, and particularly by tunnels.
So far we have only considered the theory, even more important are the practical challenges of collecting usable data on a comparable basis. In the period that you are talking about, the professional engineers could use a dynamometer car for measurement, but the amateur enthusiast only had manual stopwatches. Having also done manual timing at competitive sport, I know that there is significant measurement error and bias in manual timing.
The much more basic technology also applied to locomotive maintenance: there were no electronic engine management systems and key components such as fuel injectors and turbochargers would be adjusted by depot engineers. As a result there would be considerable variation in performance between members of the same class.
I was assuming a standard 315 TL
On most routes there was no such thing as a standard trailing load. Except where governed by platform length restrictions, fixed formations only really started with the HSTs then spread to hauled stock. To give one example, in the 1970s the 2 Cambridge Buffet Express sets were 8 cars, but every weekday afternoon 2 extra coaches would be attached to the 1530 from Cambridge, then removed again when the set got back to Cambridge on the 1714 from Kings Cross.
Another factor where there was much more variation in those days was driving technique, and this had most impact at starts and stops. Remember that trainee drivers learned their techniques from more senior colleagues, on the job, not on a simulator. In particular drivers would not all learn the same braking points.
Taking all of these together, valid comparisons required elimination of as many sources of error as possible. For this reason trains running at a steady speed on a steady gradient offered the least unreliable comparisons.