
Intelligent Test Recommender (ITR): Change-Aware, AI-Driven Test Selection in CI/CT Regression Pipelines
A Joint Technical Whitepaper — NXP Semiconductors · Soliton Technologies
Abstract
Continuous Integration and Testing (CI/CT) has become the de facto practice in most Software Development organizations. Code merge requests into a mainline code branch are thoroughly tested for regressions at various levels – unit, component, sub-system, and system level. Any regressions are quickly detected (typically within a few hours) and fixed. However, as codebases expand and grow across multiple subsystems, regression test suites also grow in parallel, resulting in a significant increase in testing time and consuming significant amounts of computing, test infrastructure, and engineering resources. Often, these regression tests are a carefully chosen list by an expert, however over a period, it becomes very difficult to dynamically maintain, keeping it up to date. As high as 50% of regression tests are wastefully run, which yield no direct value in catching regression. Successful CI/CT implementations demand a more dynamic, adaptive test selection that can find regressions fast with the least infrastructure investment.
This paper presents Intelligent Test Recommender, an Agentic AI solution that accurately interprets code changes and dynamically recommends only those tests that are impacted by it. By establishing structured relationships between the codebase, functionality, and the test suite, the system enables change-aware test recommendations that are both traceable and auditable — replacing brute-force regression with focused, evidence-based execution.
Benchmarked across a representative set of code commits against a suite of approximately 2,000 tests, the system achieves 85–90% recall while reducing projected test execution volume by more than 60% — maintaining coverage of the tests that matter while eliminating unnecessary runs.
The solution is not specific to a particular product family, rather any domain with a structured codebase and a defined test suite can adopt the same pattern, making Intelligent Test Recommender a transferable foundation for intelligent validation at scale.
1. Introduction: Regression Testing Bottleneck
Semiconductor System Software has grown in complexity especially with the advent of Software Defined products. Early integration and delivery of fully qualified system software that meets functional and performance needs is what is now demanded by customers. It enables a faster time to market. However, as use cases grow, the test suite has also grown in parallel: more subsystems, more Software variants, more operating modes, more tests. This approach is logical, however, challenging to scale.
Consider a multi-variant SoC product family developed for automotive audio and infotainment applications, where Software is organized across multiple subsystems. Regression Validation runs at three cadences: a lightweight gate at pull-request time, a broader nightly selection, and a full weekly regression of thousands of test cases executed end-to-end over the weekend. When the weekly run overruns its window, triage cycles slip and the entire release cadence absorbs the delay.
The underlying inefficiency is structural. Many changes in any given week touch only a fraction of the Software surface, yet the pipeline has no mechanism to reason about which tests are relevant to which changes — so it runs all of them. Compute, energy, and engineering hours are spent on tests that could not possibly fail given what changed, and the cost compounds with every release.
The need is therefore not a faster test runner but an adaptive one: a system that understands the relationship between code changes and test coverage and can recommend — with confidence and traceability — exactly which tests need to run, and why. Critically, this includes mapping system-level tests to Software code changes where no direct correlation exists — bridging the gap between low-level Software modifications and high-level validation coverage. This paper argues that change-aware, intelligent test selection, built on knowledge graphs and agentic AI, is the right architectural response.
2. Solution Architecture
Intelligent Test Recommender is built on two cooperating subsystems. The first consolidates discrete sources of engineering knowledge into a structure that can be queried. The second reasons over that structure at pull-request time to produce a focused test recommendation.
Subsystem 1 — Knowledge Consolidation
Software code, test case definitions, device specifications, and engineering decision records exist as distributed artifacts across repositories and document stores. A data pipeline ingests these sources and builds a knowledge graph that captures both the entities (behavioral features and code modules) and the relationships between them — including how test cases map to those entities. The graph provides relational structure and semantic discoverability that flat document stores cannot. Whenever new features or functionality are introduced, the pipeline re-runs and the graph is updated.
Subsystem 2 — Agentic Test Recommendation
At pull-request time, an agentic AI workflow consumes the code diff, traverses the knowledge graph, and reasons about which behavioral features are impacted by the change. It then maps those features to the test cases that exercise them, producing a ranked, traceable recommendation. The orchestration is built on Anthropic’ s Agent SDK, with Claude models performing reasoning steps guided by custom domain prompts.
Conceptual Flow

Figure 1. End-to-end Intelligent Test Recommender Architecture View
3. Benchmark Definition
Every AI system needs a benchmark — a shared frame of reference that defines what "good" looks like and provides a common basis for comparing iterations. Before any solution work began, we defined an internal benchmark with five components: test data, evaluation metrics, evaluation method, performance metrics, and tracking leaderboard.
Test Data
We curated a representative set of historical pull requests, each capturing the files and logic modified in the Software. For every PR, validation experts manually recorded the set of tests that should be executed. This expert-labelled set serves as ground truth against which the system is evaluated.
Evaluation Metrics
Choosing the right metrics is critical: they must reflect what experts and stakeholders consider a fair measure of system performance. We prioritized recall over precision, because in a safety-conscious validation context, a missed test is a far more serious failure than a spurious recommendation. Alongside recall, we also tracked precision, the count of unique tests recommended, and the percentage reduction in test execution volume.
Evaluation Method
We built an automated evaluation harness that takes a PR change, runs it through the Intelligent Test Recommender solution, compares the recommended set against the golden set, and provides metric values. Automation makes it efficient to re-run the evaluation as the AI system evolves through different models, prompts, and agentic architectures.
Performance Metrics
Beyond recommendation quality, we tracked end-to-end latency, token consumption per PR, and graph-build time. These operational metrics ensure the system remains practical for deployment inside a CI/CT pipeline, where wall-clock cost per recommendation directly affects adoption.
Leaderboard
An internal leaderboard records each variant of the solution — different models, agent architectures, prompts — and its metric scores. This makes it straightforward to pick winners on evidence, compare behavioral changes as the system is tuned, and avoid the trap of optimizing on intuition alone.
4. System Knowledge and Dataset Curation
The next critical component for any AI system is the knowledge it needs to make sound decisions. Rarely is all that knowledge already digitized in a form that the AI can consume. This is one of the most challenging parts of any AI initiative — ensuring enough of the relevant knowledge is captured digitally for the system to operate effectively.
In our case, the Software code repositories and the test suite were already available. Together, these provide the input context (the Software codebase) and the candidate output space (the list of tests the system can choose from). What they do not provide is device-level understanding: how features behave, what design decisions shaped the implementation, and how engineers reason about tradeoffs.
To close this gap, we curated an additional layer of device knowledge — feature descriptions, decision records extracted from technical documents, and codified engineering experience. This material gives the LLM the device-level context and decision-making heuristics it needs to interpret a code change in functional, behavioral terms.
5. System Tuning and Failure Modes
A common assumption is that most of the effort in an AI system goes into perfecting the AI itself — the model, the prompt, the agent architecture. That assumption is only partly correct. With a benchmark and leaderboard in place, iterating on AI components is relatively quick. The greater investment is in the data: understanding why the system fails on a given PR, tracing the agent's reasoning (which requires good AI observability), and identifying the gap in the underlying knowledge that produced the wrong answer. The system is, ultimately, as good as the data it is grounded in.
Once deployed, the system is monitored continuously and tuned against two distinct failure modes.
Failure Mode 1 — Incorrect Feature Extraction
The agent extracts wrong or irrelevant behavioral features from the code diff. The root cause is typically insufficient domain grounding — the model lacks the context needed to interpret what a code change means in terms of device functionality.
Resolution: Enrich the domain-specific knowledge base by adding deeper device context, communication semantics, edge-case behaviors, or clearer decision rules, until feature extraction is consistently accurate.
Failure Mode 2 — Correct Feature Extraction, Incorrect Test Recommendation
The agent extracts the right features, but the recommended tests are still wrong. The root cause lies in the knowledge graph itself — features are not correctly mapped to test cases, often because the test definitions do not expose the features they exercise in a discoverable way.
Resolution: Update the knowledge graph to correct or enrich the relationships between test cases and behavioral features, so the right tests surface for the right features.
Current state: Both tuning activities are performed by experts today. Identifying which failure mode is occurring and applying the corresponding correction requires human judgment (human-in-the-loop). Automating this feedback loop — so recommendation misses are diagnosed and used to update either the prompt context or the knowledge graph without human intervention — is a defined item on the roadmap and a step toward greater autonomy.
6. Results
The system was evaluated against the benchmark dataset, deliberately sampled across varying categories and sizes — from small, targeted bug fixes to larger changes spanning multiple functions and subsystems. This diversity was intentional, ensuring the evaluation reflected real-world variability rather than a curated best case.
The test suite used for evaluation comprised approximately 2,000 test cases. Ground truth was manually prepared for every benchmark PR, defining the exact set of tests that should be executed.
Headline Metrics
Metric | Result |
Recall | 85–90% |
Precision | 55-60% |
F1 Score | ~ 62% |
Unique test cases recommended (per PR set) | 700–800 |
Reduction in test execution volume | > 60% |
Interpreting the Results
Recall is the primary metric of confidence. At 85–90%, the system correctly identifies most tests that should run for a given code change. In a safety-conscious validation context this is the metric that matters most — a missed test is a far more serious failure than a redundant one.
Precision at 55 to 60% reflects an expected early-stage characteristic. When the system recommends more tests than necessary, the root cause is typically insufficient domain grounding — either in the prompt context or in the feature-to-test mappings in the knowledge graph. When an unexpectedly large set is recommended for a small, well-understood change, it signals that engineering judgment has not yet been fully encoded. Precision is therefore directly addressable through targeted tuning and is a defined focus area for the next phase of work.
The headline outcome is that across the benchmark PR set, the system recommends 700–800 tests out of a 2,000-case suite — a reduction in test execution volume of more than 60%, with coverage of the tests that matter preserved.
7. Deployment
Intelligent Test Recommender was introduced without modifying the existing CI/CT pipeline. The current pipeline — triggered on a nightly, weekly, and per-PR basis — drives the standard test execution workflow and was left entirely intact.
Alongside it, we established a parallel pipeline. Triggered by the same events, this parallel pipeline consumes incoming PR information and executes the full Intelligent Test Recommender workflow autonomously: retrieving the code change from version control, extracting impacted features through the Knowledge Consolidation subsystem, and producing test recommendations through the Agentic Recommender.
This parallel deployment serves three purposes:
- System monitoring — observing recommendation behavior against real PRs across varying change sizes and categories, in a live environment.
- Performance evaluation — continuously assessing recommendation quality and identifying tuning opportunities in either the domain prompt context or the knowledge graph.
- End-to-end validation — confirming the complete pipeline, from feature extraction to test recommendation, operates reliably under production-equivalent conditions.
This non-invasive deployment strategy allows recommendation quality to be measured and improved progressively, without introducing risk to the existing validation workflow. Full integration — where recommendations actively drive test execution within the production CI/CT pipeline — is the logical next step, contingent on reaching the target precision threshold.
8. Engineering Lessons
Establish the right evaluation framework early
Adopting precision and recall as the core metrics, with recall prioritized, provided a principled basis for assessing performance and directing improvement. The operational priority is clear: if recall is insufficient, focus on ensuring the system identifies the right tests before addressing over-recommendation. Sequencing improvement in this order — recall first, precision second — produces a structured path toward a steadily improving system.
Behavioral testing requires behavioral entities
A critical insight emerged from the nature of the test suite. The validation being performed is not unit-level; it is behavioral, exercising Software operation end-to-end across functional scenarios. Early iterations of the entity vocabulary were defined at a lower level of granularity than the test cases were written against, producing weak feature-to-test linkages in the knowledge graph. Reorienting the entity vocabulary to match the behavioral abstraction level of the test suite — rather than the implementation level of the Software — proved essential to graph quality and downstream recommendation accuracy. Defining entities that mirror the behavioral layer of the test suite is a foundational requirement of this approach.
Domain-specific prompt context is iterative
The domain prompt context cannot be fully specified upfront. It is developed incrementally through observed system behavior, tuning cycles, and the progressive encoding of domain expertise. Each iteration surfaces knowledge that was previously implicit — about device behavior, edge cases, and engineering judgment — making the system progressively more aligned with how experienced engineers reason about code changes and their test implications.
Separating failure modes enables targeted tuning
Recognizing that incorrect recommendations arise from two distinct root causes — incorrect feature extraction at inference time versus incorrect feature-to-test mappings in the knowledge graph — is essential for effective improvement. Without this distinction, turning effort risks being misdirected. Correctly diagnosing the active failure mode before intervening is now a standard part of the system improvement process.
It is a data problem, not an AI problem
The most consistent finding across the project was that progress was bounded far more often by gaps in the underlying knowledge than by limitations of the model or agent architecture. Once a benchmark and leaderboard are in place, iterating on AI components is comparatively fast. The slower, more valuable work is closing the data gaps that the AI exposes.
9. Limitations
The evaluation presented here is based on a benchmark defined jointly by the NXP and Soliton teams. While results demonstrate high recall and meaningful reduction in test execution volume, the evaluation scope does not yet support broad generalization of performance claims across domains or test-suite scales.
Precision at 55-60% reflects the current state of domain grounding and entity-mapping refinement — both of which are ongoing activities. Improvement in this metric is directly addressable through targeted enrichment of the domain prompt context and knowledge graph mappings and represents a defined focus area for subsequent work.
More broadly, the effectiveness of the methodology is contingent on the quality and completeness of available behavioral artifacts. In domains where interface specifications are incomplete, test cases lack sufficient interpretability, or Software behavior is not adequately documented, the methodology will encounter meaningful constraints. The degree to which behavioral-level testing can be represented in the knowledge graph determines the performance ceiling for any given deployment. Equally, the effectiveness of the agentic reasoning layer depends on the quality of context engineering — how well the prompts and contexts are structured to guide the AI in interpreting relationships, inferring relevance, and generating traceable recommendations. Poor prompt design can limit the system's ability to leverage even a well-constructed knowledge graph.
An additional constraint is the context window capacity of the underlying LLM — as the volume of code diff, graph context, and test definitions passed to the model grow, retrieval strategies and context prioritization become critical to maintaining reasoning quality.
10. Conclusion
Software validation at scale presents a fundamental tension between coverage and velocity. As SoC complexity grows and release cycles accelerate, the cost of exhaustive regression testing compounds — consuming compute, extending pipeline runtimes, and compressing the time available for engineering analysis. Expanding the test suite in lockstep with the codebase is not a sustainable trajectory.
Intelligent Test Recommender addresses this challenge by establishing a structured, query able relationship between Software changes and the test suite — one that enables change-aware, evidence-based test selection rather than brute-force execution. By grounding the system in behavioral entities derived from the device specification, and by using an agentic reasoning layer to analyze changes at inference time, the methodology delivers focused recommendations that are both traceable and auditable.
Early results demonstrate that the approach is viable. Recall of 85–90% confirms that the system reliably identifies the tests that matter, while a reduction in projected test execution volume of more than 60% demonstrates meaningful efficiency gains. Precision improvement, addressed through ongoing domain grounding and knowledge graph refinement, represents the clear path to a production-ready system.
The broader significance of this work lies in its generalizability. The methodology is not specific to any single Software domain or product family. Wherever structured specifications and behavioral test suites exist, the framework can be deployed — making Intelligent Test Recommender a transferable foundation for intelligent validation at scale across the semiconductor industry.
— NXP Semiconductors · Soliton Technologies —
Authors:
Rajkumar Anantharaman, NXP
Biju Krishnankutty, NXP
Hitesh Jaju, NXP
Aryan Kashyap, NXP
Dinesh B, NXP
Arun Natarajan, Soliton Technologies
Navin Subramani, Soliton Technologies
Govarthanam Krishnasamy, Soliton Technologies

