Industries
Home / industries / Financial Services & Banking / scenario
ILLUSTRATIVE SCENARIO
Financial Services: The Manual Creation Trap
Major Australian bank - retail and business banking | Approximately 35,000 staff | Australia
Engagement type: Unstructured Synthetic Data: GenAI Testing and Evaluation
The following scenario illustrates how organisations in this sector typically encounter the AI testing problem and how a Data Infusion engagement addresses it. It is constructed from our operational, regulatory, and technical understanding of this environment; not from a specific client engagement. It is presented as an illustrative scenario to demonstrate how the problem manifests and how it can be resolved.
The situation
The bank’s customer experience transformation programme had identified AI-assisted complaint triage as a strategic priority. Across its retail and business banking divisions, the bank received hundreds of thousands of unstructured customer contacts each month including complaints, disputes, hardship notifications, and general correspondence; routed manually through a tiered customer resolution function. The process was slow, inconsistent, and difficult to audit at scale. Escalation errors were generating downstream regulatory exposure. The volume was outpacing the capacity of the existing team.
The AI triage system had two core functions. The first was compliant classification: accurately identifying contacts that constituted formal complaints under the bank’s internal dispute resolution process and Australian Financial Complaints Authority (AFCA) referral obligations. The second was hardship identification: detecting contacts from customers who may be experiencing financial difficulty and flagging them for proactive referral to the bank’s hardship assistance team before they reached arrears, default, or AFCA.

The hardship identification function carried direct regulatory weight. Under ASIC regulatory guidance and APRA prudential expectations, banks are required to have proactive and systematic processes for identifying customers in financial difficulty. A failure to identify and assist eligible customers is not an operational inconvenience; it is a compliance failure with board-level visibility and regulatory consequence.
The development team, operating within the bank’s AI governance framework, made a deliberate decision not to use real customer correspondence for testing. The decision was correct. The approach they chose to implement it was not.
A dedicated team of eight analysts was assembled and tasked with manually constructing a test dataset: writing synthetic customer emails, complaint letters, web form submissions, and chat transcripts designed to represent the full range of contacts the triage system would encounter. The logic was sound - internal staff creation meant no privacy risk, no consent obligations, and no regulatory exposure. The bank’s second line risk function reviewed and endorsed the approach. The Chief Risk Officer was briefed.
Over five months, the team produced approximately 3,200 test documents. The dataset was reviewed by the compliance function, validated by the development team, and used as the primary evaluation corpus throughout development. The system performed strongly in testing. Complaint classification accuracy exceeded the agreed threshold. Hardship identification accuracy appeared acceptable. Deployment to a controlled pilot across two customer resolution centres was approved.
The Problem
Within eight weeks of the pilot commencing, the customer resolution leadership team identified an anomaly. The volume of proactive hardship referrals being flagged by the AI system was substantially lower than the volume being identified by experienced resolution officers working alongside it, despite handling the same incoming contacts.
An internal review was commissioned by the Chief Customer Officer. The review team extracted a sample of contacts the AI system had not flagged for hardship referral and subjected them to manual review by senior resolution officers. The finding was consistent and unambiguous: customers experiencing genuine financial difficulty were not being identified, because they were not communicating their circumstances in the explicit language the system had been produced to recognise.

The hardship identification function carried direct regulatory weight. Under ASIC regulatory guidance and APRA prudential expectations, banks are required to have proactive and systematic processes for identifying customers in financial difficulty. A failure to identify and assist eligible customers is not an operational inconvenience; it is a compliance failure with board-level visibility and regulatory consequence.
The development team, operating within the bank’s AI governance framework, made a deliberate decision not to use real customer correspondence for testing. The decision was correct. The approach they chose to implement it was not.
A dedicated team of eight analysts was assembled and tasked with manually constructing a test dataset: writing synthetic customer emails, complaint letters, web form submissions, and chat transcripts designed to represent the full range of contacts the triage system would encounter. The logic was sound - internal staff creation meant no privacy risk, no consent obligations, and no regulatory exposure. The bank’s second line risk function reviewed and endorsed the approach. The Chief Risk Officer was briefed.
Over five months, the team produced approximately 3,200 test documents. The dataset was reviewed by the compliance function, validated by the development team, and used as the primary evaluation corpus throughout development. The system performed strongly in testing. Complaint classification accuracy exceeded the agreed threshold. Hardship identification accuracy appeared acceptable. Deployment to a controlled pilot across two customer resolution centres was approved.
The Cost of the Manual Approach
This is worth being precise about, because it illustrates the compounding nature of the failure.
-
Eight analysts were allocated for five months. A material investment of skilled internal resources.
-
Salary cost, management oversight, compliance review time, and second line risk engagement across the programme.
-
Five months of development timeline during which the model was being built and calibrated against data that did not reflect its production environment.
-
Pilot deployment costs: infrastructure, change management, training, and stakeholder communication across two customer resolution centres.
-
Post-pilot remediation: model recalibration, revised deployment timeline, and an internal review process that escalated to the Chief Customer Officer, the Chief Risk Officer, and the board’s risk committee.
The cost of remediating the failure exceeded the cost of the original manual creation effort by a material margin. The five months of analyst time invested in constructing the original dataset was not recoverable. The regulatory exposure created by deploying a hardship identification system that was not performing its compliance function required formal assessment and documentation, regardless of whether external notification was ultimately required.
The Approach
A synthetic unstructured dataset of 4,800 customer correspondence documents was constructed, reflecting the bank's specific customer profile across retail and business banking, its complaint taxonomy, and its AFCA referral categories.
The engagement began with a structured scoping session involving products, policies, customer resolution, compliance, and AI development teams. The session mapped the specific linguistic patterns, contact scenarios, and edge cases relevant to complaint classification and the detection of customers experiencing financial difficulty including the bank's own internal analysis of the contact characteristics most strongly associated with financial stress in its customer base.

NUROSCEND™ constructed the dataset to reflect the full range of how real retail and business banking customers communicate including the informal register of customers writing under financial pressure, the indirect language of customers who are embarrassed about their circumstances, the fragmented correspondence of customers who are overwhelmed, and the implicit signals embedded in contacts that begin as billing enquiries and reveal hardship circumstances mid-correspondence. The dataset included representation across customer demographics, literacy levels, and communication channels, reflecting the linguistic diversity of a major bank’s actual customer population.
​
CALTREN™ validated the dataset against the bank's complaint classification framework and AFCA referral categories - both provided by the client at the scoping session, before delivery. The complete dataset was delivered with full construction methodology documentation structured for submission to the bank’s second line risk function and for inclusion in the AI programme’s formal governance record.
The Outcome
Testing against a single representative dataset is not sufficient for a complaint classification system that must reliably identify customers in financial hardship. Hardship, like fraud, is a rare event in any contact population. A model can achieve high overall accuracy simply by classifying everything as non-hardship, because most contacts are not hardship contacts. That illusion of accuracy is precisely what the original manual dataset had produced.
Data Infusion delivered three structured datasets rather than one. The first reflected the realistic contact distribution for this bank, approximately five per cent of contacts involving financial hardship signals. The second was balanced at an equal split between hardship and non-hardship contacts, deliberately loading the rare-event category to pressure-test the model's detection capability. The third inverted the distribution entirely, with hardship contacts comprising the large majority, pressure-testing the model's ability to avoid false negatives at volume.
​
A model that genuinely detects hardship performs consistently across all three datasets. A model that has learned to pass a single representative dataset fails visibly on the loaded and inverted sets because it has learned statistical shortcuts, not the linguistic signals that identify a customer in difficulty.
​
Evaluation against the three-dataset protocol identified fourteen classification failure modes not detected during the original manual data testing phase.

Nine involved correspondence from customers experiencing financial hardship expressed through indirect, colloquial, or emotionally variable language. Three involved correspondence from customers from non-English-speaking backgrounds whose phrasing patterns had not been represented in the manually constructed dataset. Two involved correspondence relating to deceased estate management and joint account disputes - categories with specific regulatory sensitivity under the bank’s AFCA obligations.
​
The model was recalibrated against the relevant failure categories before re-deployment. At the six-month post-deployment review, hardship identification accuracy had improved significantly across all previously failing categories, and proactive hardship referral volumes were consistent with the levels the original project specification had projected.
​
The second line risk function was presented with a documented, defensible testing methodology; replacing the manually constructed dataset that had carried implicit quality risk with a constructed dataset carrying explicit validation documentation. The board’s risk committee was briefed. The documentation was assessed as meeting the standard required under the bank’s AI governance framework and incorporated into the programme’s regulatory evidence file.
​
The total cost of the synthetic data engagement including the three structured datasets and the governance documentation, was less than four weeks of the salary cost the bank had already spent on manual dataset construction. The manual dataset had produced one test corpus. The Data Infusion engagement produced three, each structured to expose a different failure mode, plus a testing protocol the bank could repeat with refreshed datasets as the model and regulatory environment evolved.
What This Demonstrates
Manual test data construction is the approach large institutions reach for when they are trying to do the right thing. It is privacy-safe, well-governed, and well-intentioned. It is also expensive, slow, and produces data that does not reflect how real customers communicate under real circumstances.
The failure in this scenario was not a governance failure. The bank followed its AI governance framework. It obtained second line endorsement. It briefed the Chief Risk Officer. The compliance function reviewed the dataset. Every process was followed correctly.
The failure was a data quality failure, and it was invisible to every governance process because those processes were designed to assess whether the right steps had been taken, not whether the data produced by those steps reflected the real world. A standard accuracy measurement against a manually constructed dataset cannot reveal whether that dataset contains the linguistic characteristics the model needs to learn, because it cannot reveal what is absent.
BCG’s 2024 research found that 74% of companies are not generating tangible value from AI, and identified data governance and accessibility - not the model, as the primary reason. This scenario illustrates one version of how that plays out at the level of a single AI programme.
For a bank operating in a regulated environment, deploying an AI system that cannot recognise how customers actually communicate financial difficulty is not a marginal performance issue. It is a risk with board visibility and regulatory consequence.
​
Synthetic-but-realistic customer correspondence constructed at scale, with documented methodology and controlled linguistic variability, delivers what manual construction cannot: a test corpus that reflects the real population of contacts the AI system will encounter including the edge cases, the implicit signals, and the informal register that briefed writing never produces.

