Our new ticketing site is now live! Using either this or the original site (both powered by TrainSplit) helps support the running of the forum with every ticket purchase! Find out more and ask any questions/give us feedback in this thread!
More TOCs are making available individual train loading data including South Western, South Eastern, GTR, Northern and Merseyrail at raildata.org.uk. The data is not available for all services and there are sometimes issues with accuracy. Merseyrail has just withdrawn its data for this reason.
Access to the data is free but requires registration.
More TOCs are making available individual train loading data including South Western, South Eastern, GTR, Northern and Merseyrail at raildata.org.uk. The data is not available for all services and there are sometimes issues with accuracy. Merseyrail has just withdrawn its data for this reason.
Access to the data is free but requires registration.
Why is that the case, particularly over a reasonable time period? Neither method is likely to be fully accurate but over time one might have expected the variation in individual counts to even out a bit, or at least a consistent under/over estimation factor to be derived.
Which way are the figures out?
Or is this derived from ticketing data which relies on an algorithm to allocate passenger journeys to trains and that algorithm doesn’t reflect the number of journeys per ticket, the distribution of travel over the validity of that ticket or actual routings terribly well?
Or is it derived from mobile phone data from one provider and badly grossed up for some reason?
Still a very interesting data source which has taken a long time to become available in the public domain, however inaccurate it appears to be at the moment. This should improve in time. One benefit of ‘nationalisation’, as previously this kind of data would have been deemed commercially sensitive.
Short version: where the approach mentioned by @Clarence Yard is used it can lead to really nasty problems if you try to make decisions using the results. Probing one inconsistency gets into a cycle of worsening evident problems.
Quick thoughts (I have not looked at the data yet):
1) The distinction between observed data and model outputs is fundamental (for onward use in models and decision making) BUT seems difficult to get through to people at all levels from operation up to management (I had 40 years' experience of that from time to time). Road traffic counts are notorious too.
2) I can guess the shape of that demand forecasting model. Beware what could happen next: another model called maximum entropy could be applied to bend the first model outputs to fit observed data to 'solve' the problems people complain about. In turn that is likely to introduce worse fit in the places where a) you have no observations and b) bad observations cause wider pollution.
== Doublepost prevention - post automatically merged: ==
I would disagree. Technology and algorithms get more trust than rigorous analysis. Salesman says 'this is true', statistics says there is wide range of answers. Management goes with the salesman's view.
Why can't they get TOC loadings? Where I work, if we were aware of such data going out, we would object, but sometimes you get overruled or you don't have the authority to stop people.
I have worked during the past on a project for NR where they very much did get loading data from most TOCs - to my knowledge they still are getting that data. The project aim was to determine load on every passenger train on NR metal without using TOC based data sources, but the loading data from TOCs was used to test for validity and accuracy. That project still exists with a different supplier to this day to the best of my knowledge, and the data out of it is available internally in the industry but not as open data due to some of the sourcing - again caveated with accuracy/validity issues as with all these things.
I suspect the datasets that are released by NR are on the basis of the open data principles that NR are following at the moment - which in general is there any really compelling reason to not do so...
For SE the data is based on "loadweigh" - each carriage is weighed in real-time with this then divided by an "average" figure per passenger. This is then aggregated over 28 days to generate the average loading for each train on each day of the week.
I have worked during the past on a project for NR where they very much did get loading data from most TOCs - to my knowledge they still are getting that data. The project aim was to determine load on every passenger train on NR metal without using TOC based data sources, but the loading data from TOCs was used to test for validity and accuracy. That project still exists with a different supplier to this day to the best of my knowledge, and the data out of it is available internally in the industry but not as open data due to some of the sourcing - again caveated with accuracy/validity issues as with all these things.
I suspect the datasets that are released by NR are on the basis of the open data principles that NR are following at the moment - which in general is there any really compelling reason to not do so...
If the data is inaccurate then I would say don't publish.
Generally I would be for publication though but not if it's inaccurate. And I don't mean it needs to 100% perfect as nothing is but a level better than this.
We are aware of an issue with emails from the Forum to Microsoft-based email accounts (hotmail/outlook/live.com email addresses). This is being looked into currently, thanks for your patience meanwhile.