Summary – TLDR
- A large language model (LLM) can perform various tasks, such as generating text. It can also be multimodal (text, images, audio).
- LLMs can be deployed either in the cloud or on-premise (a server on site).
- Using commercial LLMs such as ChatGPT, Claude and Gemini can make sense, as long as their use is framed by a clear internal policy.
- It’s important to verify the information, because an LLM can hallucinate.
- The benefits for the mining industry are numerous, including:
— Optimizing operations and improving safety through data analysis. — Extracting valuable information from a variety of data sources. — Generating reports such as: safety incident analysis, training support, environmental risk management, performance optimization.
What is an LLM?
A large language model (LLM) is a type of artificial intelligence model designed for a range of tasks, from simple classification through to highly advanced text generation. Picture a tool that has read millions of books, articles and websites, and can use that knowledge to write text, answer questions or even hold a conversation. These models use sophisticated algorithms to identify patterns in language and predict the words or sentences that come next. They’re called “large” because they process enormous quantities of data and parameters to work effectively. Ultimately, an LLM can act as a very intelligent virtual assistant able to work with text in an almost human way.
While language is central to LLMs, there are also multimodal LLMs that aren’t limited to text data and can handle other types, such as images and audio.
On site or in the cloud?
Running an LLM on-premise rather than in the cloud has several important advantages, especially for companies sensitive to data security and confidentiality — a mining company, for instance.
Data security: keeping the data on site limits the risk of leaks or breaches that can occur when data moves to and from the cloud. That’s particularly crucial for companies handling sensitive information.
Confidentiality: the data stays inside the company and isn’t exposed to third parties. That helps protect intellectual property and confidential information.
Full control: the company has full control over its infrastructure and its data. It can customize and optimize the model for its specific needs without depending on a cloud service provider.
Performance: LLMs can be very resource-hungry. Depending on the size of the company, hosting them on site can minimize latency (network latency is the delay in network communication — the time needed to move data across the network) and optimize performance for the infrastructure available.
In short, having an LLM on site means better protecting and controlling the data while optimizing performance and meeting the company’s specific needs.
Yes, but what about ChatGPT, Claude and Gemini?
Using a public commercial LLM such as ChatGPT or Gemini in a company securely and without compromising data confidentiality is possible, by focusing on common, low-risk applications such as:
Writing content: these tools can help draft articles, reports, emails or social media posts. Those tasks can be done without disclosing confidential data, as long as you don’t include it while drafting.
Brainstorming and ideation: these tools can be used to generate creative ideas, suggest solutions or support brainstorming sessions, without compromising information security.
By using ChatGPT or Gemini for these kinds of applications, companies can benefit from their capabilities while maintaining a high level of data security and confidentiality.
Note that enterprise versions of these tools exist and come with more secure terms and conditions for the data sent to them. Nothing is foolproof, and the terms can change over time without that being made very explicit. So we strongly recommend staying attentive, and framing the use of these tools in the company with a clear policy.
What’s very important to understand is that the information supplied by these third parties can be wholly or partly false, and you have to be able to judge the generated content and understand what it was generated from. So ideation or drafting applications are ideal for this kind of tool in its current state.
Companies have access to different working methods and policies to govern the use of online LLMs. It’s likely these tools are already used in most companies, and in an ungoverned way, because they’re very useful and deliver big efficiency gains on certain tasks. So it’s very important to establish an internal policy and put ways in place to offer these solutions while protecting your data and your secrets.
Hallucinations?
LLM hallucination happens when the model generates incorrect or invented information that looks plausible but isn’t grounded in real data. Put another way, it’s as if the AI “imagined” answers instead of relying on facts.
A plain-language example:
Imagine you ask a friend for advice on a book, and instead of saying “I don’t know” or looking the information up, they invent details about it. The friend might say the book is about a dragon and a pirate when in reality the book has nothing to do with that. That’s what we call a “hallucination” in the context of language models.
Why does it happen?
LLMs such as ChatGPT and Gemini are trained on enormous quantities of text from the internet. They learn to predict the words that come after a given sequence of words, but they don’t really understand the world the way a human does. Sometimes, to make sure they give a complete answer, they can generate information that isn’t correct.
How do you avoid it?
To minimize the risks tied to hallucination, it’s important to verify the information the AI supplies, especially where it’s used in critical contexts or where high precision is needed. Using AI for tasks where errors matter less, such as generating ideas or assisting with writing, can also help avoid hallucination-related problems.
Open licence
An open-licence large language model offers several advantages over a commercial one. First, it allows complete transparency, letting researchers and developers understand and improve the model. Second, it encourages collaborative innovation, since the community can contribute to improving it and adapting it to a variety of specific needs. On top of that, open-licence models are often free, which lowers costs for the companies and individuals who want to use them. Another crucial advantage is the ability to deploy the model on site, giving full control over data and operations — particularly important for companies with strict confidentiality and security requirements.
The mining sector
In mining, managing data effectively is crucial to optimizing operations, improving productivity and ensuring safety. But mining companies often face challenges around collecting, analysing and exploiting data from a variety of sources, such as ERP (Enterprise Resource Planning) systems, CMMS (Computerized Maintenance Management Systems) and forms filled in by field staff. Companies that are further along in their digitization have tools for digitizing inspections and field data, such as Stelar or others.
The data collected through forms filled in by field staff is often under-used. By applying natural language processing techniques to those forms, companies can extract valuable information about working conditions, safety incidents and team performance. That makes it possible to improve operational practice and strengthen safety at mine sites. It’s then possible to cross-reference that data with ERP spending, CMMS maintenance and asset condition in the asset integrity system. The more the data generated and received by the company is digitized and stored consistently and in a structured way, the more it becomes possible to draw value from it.
Here are a few examples of reports that could be generated from such a solution, on top of the reports coming directly out of each system:
- Safety incident analysis (OHS)
From a history of incident reports (description, date, equipment involved, external conditions, photos) and the documentation of safety policies and procedures, generate a synthesis of incidents, their recurring causes and the corrective measures put in place. Based on internal policies and procedures, suggest improvements for prevention and employee training. Assess compliance with internal policies and procedures.
- Training support
From the history of work orders and reports, maintenance, equipment inventories and equipment data sheets, generate a training plan, training manuals and relevant maintenance procedures. Assist new talent in real time through a chat application targeted at specialized content.
- Managing environmental risk
From photos and field logs, estimate operational compliance against documented standards. Generate a daily synthesis of probable non-conformities for later analysis by an expert (pre-filtering and triage by severity).
- Performance optimization
To minimize unplanned stoppages, produce a synthesis of the equipment types needing more regular follow-up, based on maintenance history (work orders, reports).
On-site LLMs make it possible to query these different data sources all at once. As with the online tools, you have to stay vigilant: the work is greatly accelerated, but these tools can make mistakes and hallucinate.
Flexibility and scalability
A major advantage of using open-licence LLMs is their flexibility and their ability to evolve with the company’s changing needs. Companies can continually refine and improve their models based on new data and lessons learned, without depending on expensive proprietary solutions.
Conclusion
Using open-source language models offers enormous potential for exploiting on-site data in the mining sector. By integrating and analysing data from the ERP, the CMMS and field forms, mining companies can not only improve their operational efficiency but also uncover new opportunities for innovation and growth.








