Industries
Home / industries / retail energy / scenario
ILLUSTRATIVE SCENARIO
Retail Energy: The Build-or-Buy Decision
Engagement type: Unstructured Synthetic Data. GenAI testing and evaluation for customer contact classification - Australia
The following scenario illustrates how organisations in this sector typically encounter the AI testing problem and how a Data Infusion engagement addresses it. It is constructed from our operational, regulatory, and technical understanding of this environment; not from a specific client engagement. It is presented as an illustrative scenario to demonstrate how the problem manifests and how it can be resolved.
The situation
The retailer had been developing an AI-assisted system to classify and route incoming residential customer contacts including calls transcribed through their IVR system, web chat sessions, and email, to the appropriate team or response workflow. The classification task covered the main reasons a residential customer contacts their energy retailer: billing and account enquiries, payment extensions, direct debit failures, move-in and move-out requests, solar feed-in tariff questions, outage and supply issues, and general account management.
​
The decision not to use real customer contact data was made early and without significant debate. Residential energy customers share genuinely sensitive personal and financial information in their contacts such as account details, payment histories, employment changes, and family circumstances. Using real contact data for AI testing, even in a restricted internal environment, would have introduced privacy obligations the organisation was not prepared to accept. The decision was right.

The question was what to use instead. The technology team recognised that large language models could generate synthetic contacts and saw an opportunity to build a reusable internal capability rather than engage a third party. The rationale was sound on its face: if the organisation could generate its own test data on demand, it would have an asset it could deploy across every future AI project. The investment seemed proportionate to the scale of the AI roadmap ahead.
​
Over several months, the team built a prompt framework - scenario definitions, customer persona descriptions, contact format specifications, and connected it to a commercial large language model via API. They generated test contacts across all the main contact categories. The contacts looked professional and varied. They were reviewed internally, assessed as representative, and used as the primary test corpus throughout development.
​
The system performed well in testing. Contact classification accuracy across all categories exceeded the target threshold. The system was deployed.
The Problem
Performance in the live environment did not match performance in testing. Contacts were being misclassified and routed to the wrong queues, concentrated in categories that had shown strong results during evaluation.
​
The operations team pulled a sample of misrouted live contacts and compared them against the internally generated test contacts representing the same categories. The failure mode was immediately apparent.
​
The internally generated test contacts were fluent, specific, and well-structured. A contact about a billing dispute stated clearly that it was about a billing dispute. A contact requesting a payment extension explained the situation coherently and asked a direct question. A contact about a move-out gave the date, the new address, and a specific request. Each contact was a clear, complete expression of a single, already-understood customer need.

Real residential customer contacts are not constructed this way. Not because customers are inarticulate, but because they are navigating something unfamiliar or stressful, and they do not know what category their issue belongs to before they make contact.
​
The customer who calls because something looks wrong on their bill does not usually open with "I am calling to dispute my bill." They say: "I got my bill and it seems really high, higher than usual and I’m not sure why." The customer calling about a failed direct debit does not lead with that. They say: "I got a text about my account, something about a payment?" or "There was something on my bank app. I think it might be related to you?" The customer moving house may not use the words “move out” or “transfer” at all: "We’re moving at the end of the month, what do I need to do about the electricity?"
​
The prompt framework had generated contacts from the perspective of someone who already knew what the contact was about - a system designer writing scenarios, not a residential customer experiencing one. The gap between those two registers is where the classification system failed.
The Real Cost of Building It Internally
The misrouting and remediation were the visible costs. But the review that followed surfaced something more significant: the ongoing cost of operating a capability the organisation had not intended to build.
​
The internal platform the team had created was not a finished product. It was the first version of a non-core system that required continuous investment to remain useful:
-
The prompt framework needed updating whenever the contact taxonomy changed - new products, fee structure changes, regulatory updates affecting how certain contacts had to be handled.
-
The persona definitions needed to reflect shifts in the customer base - new demographics, changing channel behaviours, evolving expectations across voice and digital contacts.
-
Generated contacts needed reviewing and validation with each iteration. There was no independent validation layer. The team that built the generation framework was also responsible for assessing whether its output was adequate - a structural conflict that no internal process could fully resolve.
-
Every new AI use case on the organisation’s roadmap would require the framework to be extended: new scenario categories, new output formats, new validation criteria.
The team had set out to produce test data for one AI project. They had, without fully recognising it at the time, become responsible for operating a synthetic data platform, with all the maintenance, iteration, and governance overhead that platform operation requires. Engineering capacity that should have been directed at the AI roadmap was being absorbed by a non-core capability.
​
The decision had not been whether to purchase synthetic test data once. It had been whether to become a synthetic data platform operator in addition to being an energy retailer. Those are different decisions, with different cost structures, and the second was never the one that was made.
The Approach
The Data Infusion engagement began with a structured scoping session with the retailer’s customer operations, digital channels, and product teams. No real customer contacts were accessed at any stage. The session focused on building a detailed picture of the contact taxonomy and critically, the population of customers who experience each scenario.
​
The distinction between scenario and population is where an internally built prompt framework falls short, and where a purpose-built capability begins. A scenario is what is happening. A population is who experiences it, across the full range of customer types, emotional states, linguistic registers, and levels of familiarity with energy retail processes.
​
A customer wanting to transfer their account because they are moving is a single scenario. The population who experiences it includes: the first-time renter managing a utility account for the first time; the homeowner who has moved several times and knows exactly what to do; the elderly customer whose adult child usually handles these calls; the customer whose first language is not English; the customer who has left this to the last minute and is already stressed about the move. Same scenario. Different people. Different language. Different signals.

The scoping session mapped the contact taxonomy across all primary residential contact reasons and documented the customer population range for each category and not from the customer data itself, but from the operational knowledge held by the customer operations team. This persona framework was the foundation of the dataset construction.
​
NUROSCEND™ constructed a dataset of residential customer contact transcripts and messages that reflected this full population range across each contact category - generating contacts from the position of each customer type in each scenario, not from the position of a designer describing the scenario. The billing enquiry from a customer who knows exactly what they are disputing, and the one from a customer who only knows something seems wrong. The move-out contact that uses process language, and the one that does not use process language at all. The direct debit contact from a customer who is apologetic and embarrassed, and the one from a customer who is certain it is the retailer’s fault.
​
The constructed dataset was placed into CALTREN™, the controlled, secure environment in which Data Infusion’s synthetic datasets are placed, iterated, and validated, where it was validated against the retailer’s contact classification framework before delivery. The dataset of 3,200 contacts across the full residential contact taxonomy was delivered with full construction methodology documentation and a mapping of the persona types and linguistic registers represented across each category. No ongoing maintenance obligation transferred to the retailer.
The Outcome
Evaluation against the synthetic dataset identified the classification failures in detail. The model was re-tested using the dataset as the evaluation and testing corpus. Post-testing, classification accuracy across all contact categories met or exceeded the target threshold including indirect, hesitant, and emotionally variable language patterns that had caused the original deployment failures.
​
The system was redeployed. At the three-month post-redeployment review, first-contact resolution rates were above the system’s original targets and misrouting rates had reduced to within acceptable operational tolerance.

The technology team also closed out the internal capability they had built. The prompt framework, the persona definitions, the generation scripts - these were decommissioned. Engineering capacity returned to the AI use cases on the organisation’s roadmap rather than to maintaining a synthetic data platform it had not planned to build.
​
The customer operations and technology teams documented the Data Infusion methodology as part of the organisation’s AI governance record, covering the construction approach, the persona framework, the contact taxonomy mapping, and the validation process. The documentation provided an auditable record of how the system had been tested, produced independently of the team responsible for building it.
What This Demonstrates
The decision to build an internal synthetic data capability is not a decision about a tool. It is a decision about whether to become a synthetic data platform operator with the ongoing investment in maintenance, iteration, independent validation, and extension that platform operation requires.
​
A retail energy organisation’s competitive advantage is not in its ability to generate synthetic test data. The same logic that leads an organisation to buy a CRM rather than build one applies here. The question is not whether it could be built internally. It could. The question is whether that is where time, capital, and engineering capacity should go.
​
A purpose-built, independent testing capability delivers what an internally built prompt framework cannot: population-level realism across the full range of customers who will actually contact the system; independent validation separate from the team that built the generation process; and a governance record that is defensible precisely because it was produced by a party without a stake in the outcome.
​
The cost of the internal build in this scenario was not the hours spent writing prompts. It was the months of engineering capacity directed away from the AI roadmap, the structural absence of independent validation, and the ongoing overhead of a non-core platform that was never going to be decommissioned until a better alternative was found.

