• Our new ticketing site is now live! Using either this or the original site (both powered by TrainSplit) helps support the running of the forum with every ticket purchase! Find out more and ask any questions/give us feedback in this thread!

How trustworthy is the output from modern LLMs?

Status
Not open for further replies.

Crithylum

Member
Joined
21 May 2024
Messages
349
Location
London Borough of Ealing
Inspiration from this thread: https://www.railforums.co.uk/threads/refund-refused.301783/

(3 years ago), I would barely trust anything output from an LLM (large language model). When Gemini (an LLM produced by Google) was first added to search results, it was terrible, but now it is actually quite good. I think it is at the point now where (on average) it gives better information than if you clicked the first link (which is often sponsored). If you are willing to put the effort in, you will be able to find more accurate information, however, alot of people seem incapable of this nowadays (maybe as a result of these LLMs?).

The reason I am posting this thread is that in the linked thread, the OP (original poster) includes screenshots of information from an AI (artificial intelligence), which receives criticism in a later post, with the OP then becoming slightly defensive about trusting the AI. However, the AI summary appears to be more accurate than what the OP understands about the situation ("delays" vs altered timetable). Should we maybe start trusting AI a bit more (especially compared to the rapidly declining quality of output from "news" organisations)?
 
Last edited:
Sponsor Post - registered members do not see these adverts; click here to register, or click here to log in
R

RailUK Forums

AlterEgo

Verified Rep - Wingin' It! Paul Lucas
Joined
30 Dec 2008
Messages
29,544
Location
LBK
In the example given the LLM has answered the question put to it. The OP is the one who has not understood the wider context and their reliance on the AI response in and of itself was their undoing. Putting the question to real people has generated a much more useful response because those people are able to interrogate the OP and understand the context of the question and the motive behind it.

The OP's the one who bought the ticket and relied on something they heard on the radio to decide simply on hearing that and then asking AI a contextless question to cancel a journey they had no doubt actually booked - one the same with diversions etc.
 

JohnMcL7

Member
Joined
18 Apr 2018
Messages
1,055
Should we maybe start trusting AI a bit more (especially compared to the rapidly declining quality of output from "news" organisations)?
Absolutely not, LLMs getting it wrong is an inherent part of the way they work so anything they say should be treated as potentially wrong. The problem now is identifying what is actually LLM generated, many of the search results are now pointing to LLM generated content and people are posting forum responses which are LLM generated but not labelled as such. I see so many unbelievably stupid answers everywhere now and what I really don't understand is no matter how many obviously wrong answers these tools generate, it doesn't seem to stop people using them so I'm completely baffled why you'd think they should be trusted more. Scepticism is absolutely the correct way to treat the information from an LLM tool even if it's ultimately correct.

The type of data also can significantly vary the accuracy as well they're more likely to get static information correct whereas they come up with some really wrong answers for train timetables where they've incorrectly pieced together changing information.

The term 'hallucination' is such a brilliant piece of marketing to try and conceal how wrong these tools can be.

LLM - Large Language Model, this is the type of generative AI most of the big tools are using (like ChatGPT) where it's trained on existing data then uses probability to piece together text (a gross simplification I know). Because there is no actual intelligence there and it's using probability to link the text it can get it completely wrong with no way to verify itself.
 
Last edited:

The Pelican

Member
Joined
3 Oct 2025
Messages
424
Location
London
Post 2025 LLMs often go through a late "Reinforcement Learning" phase, where the model's answers in some objective areas, like maths and coding, are compared to the true values. This has made the models, which had previously made maths mistakes in basic sums, more reliable in these areas. I don't know if any of the companies trains their models in transport timetables. Some models are capable of accessing webpages, such as a transport companies', via something called an API, which is a method for computer programs to access a website. These might be able to run a query for you like you could run yourself at a transport planning website.
 

Bletchleyite

Veteran Member
Joined
20 Oct 2014
Messages
113,712
Location
"Marston Vale mafia"
I don't know either. Is it a station code?

It's a large language model - basically a form of AI - so called because it is "trained" on a huge body of textual data from e.g. the Internet which allows it to create a statistical model which determines what word should come next. The outcome of this is a machine that speaks as if it was human (ish), and like a human it can make mistakes and obfuscate, though whether it can lie is for debate as it doesn't have consciousness or free will as such (though may appear to).
 

deltic

Established Member
Joined
8 Feb 2010
Messages
3,719
Depends on what you want it to do and what you ask it. If you want it to summarise a document or undertake detailed analysis of a large dataset I would say it is pretty reliable. If you then want it to explain the implications of that analysis it is pretty dodgy and you can see the biases in what it comes out with.
 

styles

Established Member
Joined
7 Dec 2014
Messages
4,930
Location
Gwynedd
It really depends on the context.

I'm now using LLMs and other AI agents at my work and they honestly provide exceptional value for money. However, they are not infallible and for many situations they need a human review.

Trainline had an "AI" chatbot which gave horrendous ticketing advice. Whether it was trained on bad data, not enough data or something else, who knows, but the advice was bad. They switched it off. I mean they had to, because realistically under consumer rights law, you'd be entitled to rely on such information, even if it were wrong.

There are times and places for it, but we shouldn't use it without recognising the limitations.
 

Bletchleyite

Veteran Member
Joined
20 Oct 2014
Messages
113,712
Location
"Marston Vale mafia"
Trainline had an "AI" chatbot which gave horrendous ticketing advice. Whether it was trained on bad data, not enough data or something else, who knows, but the advice was bad. They switched it off. I mean they had to, because realistically under consumer rights law, you'd be entitled to rely on such information, even if it were wrong.

I think it's been realised that using the LLM "knowledge" to do stuff like public transport times/fares tends to produce inferred rubbish like there's a bus from A-B and C-D hourly so there would be B-C too. That being the case, good implementations are set up to use the LLM to interpret the request, research it via appropriate sources (and not make anything up that is not in those sources) and then use the LLM again to produce a human readable output.
 

bleeder4

Member
Joined
19 Jan 2019
Messages
792
Location
Worcester
LLMs are very useful for researching stuff and presenting you with a shortlist for you to investigate. For example - "I am looking for a hotel in Manchester. It needs to be within a 10 minute walk of Piccadilly station, it needs to have step-free disabled access, it needs to be dog friendly, it needs to have a family room and the hotel restaurant needs to have a gluten free option and 2 red wines on the menu. For each hotel you find that meets these criteria, analyse the online reviews and present a bullet point summary of what people liked and what they didn't like."

A prompt like that can save you a whole heap of time.
 

styles

Established Member
Joined
7 Dec 2014
Messages
4,930
Location
Gwynedd
LLMs are very useful for researching stuff and presenting you with a shortlist for you to investigate. For example - "I am looking for a hotel in Manchester. It needs to be within a 10 minute walk of Piccadilly station, it needs to have step-free disabled access, it needs to be dog friendly, it needs to have a family room and the hotel restaurant needs to have a gluten free option and 2 red wines on the menu. For each hotel you find that meets these criteria, analyse the online reviews and present a bullet point summary of what people liked and what they didn't like."

A prompt like that can save you a whole heap of time.
I'm all for AI, but you've been able to do this sort of query on hotels.com for a decade or more with tickboxes. I say this as somebody regularly travelling with a wheelchair user and a dog (albeit not usually children).
 

Tetchytyke

Veteran Member
Joined
12 Sep 2013
Messages
17,623
Location
Isle of Man
My experience of the three main LLMs- ChatGPT, CoPilot, and Gemini- is that if they don’t know the answer they will simply make stuff up.

We have a big push at work to use CoPilot more but, for the type of OSINT analysis and research I conduct in my job, it is fundamentally useless. If I have to manually check all its outputs to make sure it isn’t just making stuff up then I may as well just do the research myself in the first place.
 

Bletchleyite

Veteran Member
Joined
20 Oct 2014
Messages
113,712
Location
"Marston Vale mafia"
If I have to manually check all its outputs to make sure it isn’t just making stuff up then I may as well just do the research myself in the first place.

I don't think that's really true though. You can ask it to give its references and so you've got a good set of pointers to where to look. That still saves time even if you're having to read those sources yourself.
 

Tetchytyke

Veteran Member
Joined
12 Sep 2013
Messages
17,623
Location
Isle of Man
You can ask it to give its references and so you've got a good set of pointers to where to look.
That assumes that the references it gives actually relate to the subject matter at hand. I have seen countless examples with all three LLMs where the references don’t actually say what the LLM reports that they say.

It’s sometimes pretty basic stuff. Gemini swore blind to me that a specific company in another jurisdiction was registered at a specific address and told me it was in the companies register for that jurisdiction. It even gave me the link to prove it. Sadly, the link referred to another company in another jurisdiction which had nothing to do with the question I asked…
 

DelW

Established Member
Joined
15 Jan 2015
Messages
6,024
My experience of the three main LLMs- ChatGPT, CoPilot, and Gemini- is that if they don’t know the answer they will simply make stuff up.
They've learned to replicate human behaviour pretty accurately then :s
 

Dave W

Member
Joined
27 Sep 2019
Messages
893
Location
North London
My experience of the three main LLMs- ChatGPT, CoPilot, and Gemini- is that if they don’t know the answer they will simply make stuff up.

We have a big push at work to use CoPilot more but, for the type of OSINT analysis and research I conduct in my job, it is fundamentally useless. If I have to manually check all its outputs to make sure it isn’t just making stuff up then I may as well just do the research myself in the first place.

That assumes that the references it gives actually relate to the subject matter at hand. I have seen countless examples with all three LLMs where the references don’t actually say what the LLM reports that they say.

It’s sometimes pretty basic stuff. Gemini swore blind to me that a specific company in another jurisdiction was registered at a specific address and told me it was in the companies register for that jurisdiction. It even gave me the link to prove it. Sadly, the link referred to another company in another jurisdiction which had nothing to do with the question I asked…
"Hallucination" is a massive problem in the LLM generative AI space. Where Copilot is best deployed is not asking it questions at large, but to ask it to interrogate specific and limited datasets and apply its LLM to that. The problem that big organisations - of all types - face is that whether we like it or not, people are using it. At my place (circa 2200 people, of which 1400 have regular access to work-issued IT) 1200 separate accounts accessed Copilot in February. That's more than used PowerPoint and it was barely beaten by Excel. So the task is to ensure it is being used appropriately. Most people just want a summary of their meeting notes or emails, which Copilot is very good at.

It is less good when you ask it a question like you would a search engine, or more importantly when you get into the realm of specialist tasks (which it seems you allude to in your first post). There are many platforms being developed for specific uses like that across the piece, but with that specialism and complexity comes added cost.
 

Yew

Established Member
Joined
12 Mar 2011
Messages
7,244
Location
UK
I'm all for AI, but you've been able to do this sort of query on hotels.com for a decade or more with tickboxes. I say this as somebody regularly travelling with a wheelchair user and a dog (albeit not usually children).
That's maybe not a great example, but "I'm going to city X, i've done Y and Z attractions so far what others should I consider" can be a useful question.
 

Bletchleyite

Veteran Member
Joined
20 Oct 2014
Messages
113,712
Location
"Marston Vale mafia"
It is less good when you ask it a question like you would a search engine, or more importantly when you get into the realm of specialist tasks (which it seems you allude to in your first post). There are many platforms being developed for specific uses like that across the piece, but with that specialism and complexity comes added cost.

Though you certainly can usefully use it as a "super search engine" - ask it to provide a few summary bullet points about the thing you're searching and a list of links about it with a one sentence summary for each.

I think the difference from a search engine is that a couple of words won't generally give you good results, you do need to tell it what to do.
 

Dave W

Member
Joined
27 Sep 2019
Messages
893
Location
North London
In the spirit of holidays, I'm going away on Friday, and I've just asked my (work!) Copilot for some tips. With proper prompting (which is the key here), it's generated me a full 7 day itinerary around my two year old. I gave it options for hiring a car on one or two of the days, whilst maximising our all inclusive value and accounting for the fact she'll want an afternoon nap. I could - of course - have done this myself, but it has the potential to remove a lot of stress and anxiety. It also acts as a search engine in this context because it provides options I may not have thought about.

I think we are a long way on from the first LLMs to hit the market. And by design they're learning all the time.

The issue comes - per the OP - when it is accepted as fact; businesses - and TOCs aren't immune - will need to adapt.
 

Tetchytyke

Veteran Member
Joined
12 Sep 2013
Messages
17,623
Location
Isle of Man
It is less good when you ask it a question like you would a search engine, or more importantly when you get into the realm of specialist tasks (which it seems you allude to in your first post).
I think what I find bizarre about it is that Gemini's output is markedly worse than what I get from simply using Google's (or DuckDuckGo's) non-AI search and applying the standard operator functions, e.g. by site or by file type. I know an LLM is not a search engine and can't be used in the same way, but still.

The hallucinations are also a real problem.

The problem that big organisations - of all types - face is that whether we like it or not, people are using it.
I think it depends what people are using it for.

I use it for things such as stylistic suggestions, e.g. when I'm about to send an email which I know won't be well received by the recipient, and I want a second opinion on how to soften it. It's actually pretty good at that, within certain parameters anyway. Forum posts (particularly elsewhere, but sometimes here in the disputes section) which are obviously written by an LLM do my head in; if I read one more "key takeaway" I think I shall scream. And yes, I am grumpy that the em dash- something I've always used- is now taken as an LLM tell.

My issue with the hallucinations is that the LLM always repeats the bunkum with complete certainty. I don't blame people for taking the output as largely correct, LLM outputs don't do uncertainty and don't do nuance. If an LLM says white is black, and provides a reference, I don't blame people for not digging deeper.
 

Bletchleyite

Veteran Member
Joined
20 Oct 2014
Messages
113,712
Location
"Marston Vale mafia"
Thought I'd ask it to summarise me on here (using Grok). A good example of a hallucination came up (though most of it was unsurprising and largely correct):

(he runs a long-running thread called Bletchleyite's Bletchley musings for local observations and photos)

I did post such a thread once, but it's hardly long running, it contains 14 posts which were made over three days in summer 2018 and is now locked.

It's also interesting that it has decided that I am male. That's correct as it happens but it will have inferred that from all manner of stuff as the Forum doesn't officially state that.
 

Dave W

Member
Joined
27 Sep 2019
Messages
893
Location
North London
I am grumpy that the em dash- something I've always used- is now taken as an LLM tell.

Me too! I LOVE a dash in writing - now I've had to rein it in. In fairness if I've had copilot write me something, I then tell it to get rid of all the dashes which it does well now (in the past it just deleted them all without compunction...!)
 

bleeder4

Member
Joined
19 Jan 2019
Messages
792
Location
Worcester
Yes, I always get it to remove em dashes too. Full stops at the end of headings is another one. I always get it to remove them. Wikipedia have been listing all the tell-tale signs of AI copy as a reference for their editors, and it's now become one of the most comprehensive lists I know of. So sometimes when I'm asking ChatGPT for copy I get it to read the Wikipedia page and make sure it DOESN'T do anything on that page.


This is a list of writing and formatting conventions typical of AI chatbots such as ChatGPT, with real examples taken from Wikipedia articles, drafts, comments, and other content. It is a field guide to help detect undisclosed AI-generated content on Wikipedia: while some of the signs may be broadly applicable, some may not apply in a non-Wikipedia context.[a] Not all text featuring these indicators is AI-generated, as the large language models that power AI chatbots are trained on human writing, including Wikipedia. Many elements of AI writing can be found in editorials, blogs, or fan fiction.
 

AlterEgo

Verified Rep - Wingin' It! Paul Lucas
Joined
30 Dec 2008
Messages
29,544
Location
LBK
The output from Gemini is extremely sketchy. It hallucinates (lies) often. I asked it the same query twice:

Give me a list of all UK three letter railway station codes alongside the full name of the station, but only for codes where the first letter is M and the last letter is B.

On the first occasion it gave me only these:

Station Code Station Full Name
MDB Middlesbrough
MEB Meols Cop
MLB Millbrook (Bedfordshire)
MNB Manningtree

I asked it the same question again immediately after and it game me an accurate answer. (Where has it got MDB, MNB and MEB from, I wonder? All false)
 

gg1

Established Member
Joined
2 Jun 2011
Messages
2,498
Location
Birmingham
I use AI at work a fair bit when I'm having problems with Python code, it has it's uses but can't be relied upon.

Where the problem is a rookie mistake like incorrect syntax, AI is incredibly useful and generally reliable, for complex problems it's far more hit and miss. As an example, if I have code which works to a point but doesn't quite give me the output I'd like to see, it's not uncommon for AI to suggest amended code which either gives exactly the same result, doesn't work at all, or churns out nonsense data. If you persist with further prompts it does sometimes get there in the end but on a number of occasions it's actually taken me full circle and suggested the original code I included in the initial prompt.

The other big downside of AI when coding is the deskilling effect. I learnt SQL before AI was a thing, problems were resolved by googling and adapting someone else's solution to a similar problem to fit your own code. The advantage of doing it that way was it was a learning experience, the process of adapting the code and integrating it with mine meant the logic behind it was absorbed and I was likely to remember it if I encountered a similar problem a few months later. By contrast I'm relatively new to Python so AI has always been an option for me, when it works correctly, you literally just copy and paste the solution. It's much faster in the short term but at the cost of not upskilling you to the same degree.
 

gswindale

Member
Joined
1 Jun 2010
Messages
1,028
I'm currently working on a project to try and update a 10+ year old Excel workbook that, whilst it works fine, does rely on a fair bit of VBA code to apply protection and create different versions of the output document. Unfortunately, we now have a number of users who, for whatever reason, access the template document from within Teams and then end up trying to use it either in Teams itself or on Excel for the web.

I'm trying therefore to use Copilot to give me Office Script coding for various functions, but not being all that familiar with Javascript/Typescript, I don't always understand what the code is doing and frequently it simply doesn't work and the answer from Copilot is "you're absolutely right that this won't work as intended, so let's try something else" which then also doesn't work which results in Copilot suggesting the original code again.

There seem to be less guides around online about using Office Scripts than there were for VBA - although I do still have a copy of "Excel 2000 VBA Programmers Reference" on the bookcase, and I can't seem to find anything similar for Office Scripts as I find printed reference guides quite helpful.
 

gg1

Established Member
Joined
2 Jun 2011
Messages
2,498
Location
Birmingham
I'm currently working on a project to try and update a 10+ year old Excel workbook that, whilst it works fine, does rely on a fair bit of VBA code to apply protection and create different versions of the output document. Unfortunately, we now have a number of users who, for whatever reason, access the template document from within Teams and then end up trying to use it either in Teams itself or on Excel for the web.

I'm trying therefore to use Copilot to give me Office Script coding for various functions, but not being all that familiar with Javascript/Typescript, I don't always understand what the code is doing and frequently it simply doesn't work and the answer from Copilot is "you're absolutely right that this won't work as intended, so let's try something else" which then also doesn't work which results in Copilot suggesting the original code again.

There seem to be less guides around online about using Office Scripts than there were for VBA - although I do still have a copy of "Excel 2000 VBA Programmers Reference" on the bookcase, and I can't seem to find anything similar for Office Scripts as I find printed reference guides quite helpful.
One area I've found AI can actually be good for is explaining what a block of code does in plain English, especially useful when you've inherited someone else's code and need to troubleshoot.
 
Status
Not open for further replies.

Top