New MCP Server: Your Agents Get a Badge, Not a Master KeyMCP Server: Your Agents Get a Badge, Not a Master KeyMCP Server: Your Agents Get a Badge, Not a Master KeyMCP Server: Your Agents Get a Badge, Not a Master KeyMCP Server: Your Agents Get a Badge, Not a Master KeyMCP Server: Your Agents Get a Badge, Not a Master KeyMCP Server: Your Agents Get a Badge, Not a Master Key
Read More
Set up for scale or introducing new risks?
Michael Shearer, A former Managing Director
within HSBC Financial Crime Threat Mitigation

Agentic AI has extraordinary abilities to analyze language, make connections, and compose answers to specific questions. It can also summarize presented facts, generating accessible natural language explanations for observed behaviour.
These analytic steps of digesting and structuring information and then succinctly expressing an initial hypothesis for assessment are the backbone of the investigative process.
Agentic AI has the potential to accelerate these processes, pre-assembling a set of connected facts, and offering preliminary explanations, accelerating and complementing human judgement and critical thinking. Investigations become faster, deeper, and more comprehensively documented.
To perform these steps on behalf of an investigator, an agentic solution needs the ability to iteratively draw on diverse data sources, building out from initial leads and previously retrieved information to formulate a more complete understanding of the matter at hand. Investigations are unpredictable and often take unexpected turns as new facts emerge, requiring the scope to be adjusted and new questions to be asked of new sources.
Assembling these complex investigative jigsaw puzzles piece by piece through repeatedly querying raw data sources may ultimately be satisfactory, however it has the potential to be uncontrollably expensive and introduces new risks to the investigative process.
In this paper we will explore five reasons why disparate data blocks AI investigation agents.
Raw data sources are not typically designed for incrementally iterative data retrieval. Often, due to their organic history and legacy technology, their data has a rigid dimensional structure, tightly coupled to their primary business purpose. This increases the risk of partial retrieval as relationships between relevant data elements are overlooked, and crucial information is never surfaced for the agentic tool to process.
Consider an AI Agent performing an investigation into ‘ACME Corp’. Using an integrated tool the agent queries the financial institution’s Customer Due Diligence system for information related to companies with the name ‘ACME Corp’. The KYC file for the company is retrieved. Within the file several addresses are recorded as being associated with the company, one of which is 100 Innovation Way.
Elsewhere within the CDD system, the KYC record of a seemingly unrelated firm Example LLC also references 100 Innovation Way. However this information is not retrieved by the initial name-only query and the potential link between these organizations is missed.
Whilst an agentic process could be instructed to execute follow-on queries to search for onward connections from addresses, or many other linking attributes, this relies on the agent recognizing and extracting such entities and correlating the responses.
Relying on an agentic process to recursively retrieve and correlate all shared attributes to assemble the facts for the investigator is highly inefficient, incurring token cost for every query executed. It also increases the model management burden to demonstrate that the process is sufficiently reliable on an ongoing basis.
Delegating this process of assembling the relevant facts to a deterministic knowledge layer equips the agent with a trusted source of connected, resolved and reliable information in an efficient single query.

By Michael Shearer
Former Managing Director, Financial Crime Threat Mitigation at HSBC
Connections between data sources are challenging for an agentic data retrieval process to identify. Without understanding the situational meaning of various business terms, silos remain unconnected or worse still data is misinterpreted and assigned a meaning and significance that is not warranted. Even the smartest frontier models cannot be relied upon to correctly understand firm-specific naming conventions, burdening the investigator with checking its interpretation and spotting potential gaps.
The complexity of modern financial transactions can frustrate investigators, both human and agents alike. For example, a typical source of confusion is the reversal of actors between push and pull transactions (e.g. ACH Debit), when the originator of the payment instruction may be the payee and not the payor. The use of multi-tiered intermediary services and a divergence in transaction terminology across payment rails adds further opaqueness to the nature of transactions and the ultimate parties involved.
Human investigators are trained to spot these inconsistencies but without a clear context, the agentic process may misinterpret these transfers drawing conclusions based on inaccurate foundations.
For the agentic process to interpret these events accurately requires a knowledge layer that presents consistent, unambiguous transaction attributes regardless of nuances of individual underlying payment rails.

The risk of AI hallucinations increases in the absence of pre-connected, verified data and associations. When generative models are presented with explicitly connected matters, they are much less likely to introduce new, unsourced information into their reasoning and explanations. Language models display a propensity to fill gaps in their narratives with point-in-time knowledge from their own training process. For investigations that have similarities with firms or organizations that have a strong open-source presence, there is an increased risk of conflation leading to plausible but incorrect analysis.
Consider a financial institution investigating a new counterparty, National Rents LLC, who are conducting a high value transaction with one of their customers in a high financial crime risk jurisdiction. An agentic query on the counterparty name retrieves no confirmed exact matches from across its data repositories.
In the absence of an unambiguous negative outcome from a trusted source, the agent draws on its own training material mistakenly associating the transaction counterparty with the vehicle rental firm National Car Rental. This error leads to a narrative with misplaced confidence in the compliance regime and financial crime risk associated with this new entity.
The agentic process needs a confirmed negative result returned from a reliable, trusted knowledge layer to prevent any inappropriate recourse to its training data.

Learn how Ally applied graph analytics and contextual investigation tools to uncover complex fraud networks and strengthen fraud prevention.
Read Case StudyEmpowering an investigative agent with the unrestricted ability to query multiple data sources across an enterprise presents significant infosec and operational risks.
Operational data sources are typically subject to strict service level agreements that system owners must be confident these will be respected by any agentic query process. The burden on these platforms must be reliably managed and operational time windows respected.
From an information security perspective the task of securing and maintaining compliant, least-privilege access to data sources is complex and time-consuming.
As investigations often need to reach out to a wide variety of source systems, many of whom were not built with such ad-hoc querying in mind, the task of enabling unpredictable agentic access to all the required repositories is a significant challenge. If data sources can be connected and resolved into a single coherent source of knowledge, the problem of securely managing a single agentic interface becomes more manageable. New data sources can be integrated into the agentic process more rapidly, reusing the established knowledge interfaces and security controls.
This centralisation has the added advantage of simplifying the compliance process as the audit trail of which knowledge was retrieved and for what purpose, is readily available as opposed to piecing together an agentic story to justify queries that appear disconnected to a given investigation.
The operational impact and associated cost of repeatedly querying multiple raw data sources to assemble the investigative picture shouldn't be overlooked. Iteratively retrieving, assessing and frequently discarding irrelevant fragments of data is an expensive, time-consuming and inherently unreliable process. Agentic processes, whilst powerful and flexible, are inherently slow and prone to unexpected errors. The likelihood of getting locked in a failing query loop, racking up significant uncontrolled token costs, is a genuine risk that needs to be managed.
As new information emerges in an investigation, and the circle of subjects of interest expands, the challenge of cross-referencing new information against what is already known increases exponentially. Without clear boundaries, using an unassisted agentic process to assemble the case file can result in repeated unnecessary queries, pursuing tangential and irrelevant connections.
For example, consider an unbounded agentic investigation into an international Holding company with hundreds of subsidiaries. An agent with generic instructions to apply an exhaustive investigative discovery process to each subordinate entity could result in significant query overhead, spiraling token processing costs and an output of little investigative value.
The agentic process needs a knowledge layer that can process complex discovery queries efficiently , performing the bulk data processing once and presenting only the connected facts for consideration rather than using the model to exhaustively reason its way through the data one query at a time.

So, whilst introducing an agentic layer on top of existing data infrastructure can seem expedient, this approach increases the risk of poor investigative casework that requires expensive manual rework or manifests as latent risk to be uncovered in subsequent investigations.
Instead of assembling the investigative jigsaw piece by piece using the expensive resource of the most capable models and human analysts, an alternative approach is to delegate the task of pre-assembling the main sections of the puzzle to a unified knowledge layer equipped for the job. The agentic process can then apply its language and reasoning skills to connecting these pre-prepared knowledge fragments into a coherent hypothesis for consideration.
An Enterprise Knowledge Graph maintains that map over time, resolving identity, governing meaning, preserving provenance and enforcing who may access what, so the agent does not have to rediscover or reinterpret the business on every call.
Put differently, an enterprise knowledge graph does more than supply facts. It gives the agent a governed map of the entities, relationships and business meaning it should follow, rather than leaving the model to infer that structure for itself.
Just as data warehouses aggregated structured, numerical data for business intelligence tools to consume, now knowledge graph technology brings together rich diverse data sources, connecting information into a meaningful context that can be slotted together by the agent and critically assessed by a skilled investigator.
Consider an agentic investigative process assembling information on Acme Corp following accusations of financial misconduct.
The agent tasks the knowledge layer to identify the owners of Acme Corp and its high value transaction counterparties. The result returns a manageable number of entities, so the knowledge graph is asked to build on its initial query to identify any follow-on transactions that occurred immediately after these initial transfers. The agent then instructs the knowledge graph to find any links between any of these transacting counterparties and the Acme Corp owners to identify potential embezzlement.
The knowledge graph resolves the entities present and referring to associations documented in the KYC files identifies that the one of Acme Corp’s owners is likely a close family member of one of the transaction beneficiaries.
"Instead of 'surmising’, the agentic process uses the persistent knowledge available to process the discovered relationship between the owner and a downstream beneficiary, providing references to the transaction records and the KYC information for the investigator to verify.
The combination of agentic AI supported by a rich knowledge graph delivers an efficient and effective discovery process, accelerating and deepening investigative insights whilst controlling costs and maintaining transparency and accountability.

Agentic technology promises to accelerate the core data handling processes that lie at the heart of modern investigations. Their ability to summarise and reason has the potential to offload much of the explorative work leaving skilled investigators free to focus on assessment, judgement and critical thinking.
However to achieve their potential these agents need to draw upon relevant context-aware knowledge, assembled in a reliable, efficient and effective manner.
As you embark on an agentic journey consider the following:
These dimensions should inform your agentic knowledge strategy and how to accelerate and streamline your investigative processes.

