ChannelLife US - Industry insider news for technology resellers
United States
Monitaur launches standalone AI testing tool FlightSim

Monitaur launches standalone AI testing tool FlightSim

Mon, 28th Sep 2026 (Today)
Joseph Gabriel Lagonsin
JOSEPH GABRIEL LAGONSIN News Editor

Monitaur has launched standalone access to FlightSim, its AI governance testing product, opening the tool to companies without a full Monitaur platform licence.

FlightSim is designed to test whether AI systems behave within their intended purpose before organisations deploy them in consequential settings. It assesses systems from the outside, without requiring access to source code or model internals.

That approach is particularly relevant for buyers and operators of third-party models, where access to underlying algorithms is often restricted. The product uses repeatable statistical tests, stress tests and scenario-based checks against a system's stated requirements and intended use.

The output is a scorecard covering reliability, performance, bias and security. It can also identify issues, suggest remediation steps and provide a measurable view of how an AI system performs across risk dimensions.

Growing demand

The launch comes as businesses face growing pressure to show that AI systems can be tested consistently before being used in sensitive tasks. Some of Monitaur's large enterprise and regulated customers already require successful FlightSim results before purchasing or deploying high-impact AI systems.

This reflects a broader market shift as AI moves from experimentation into operational use. Companies in regulated sectors, or those exposed to legal and reputational risk, are looking for ways to document how systems were evaluated before approval.

Monitaur cited survey findings from the Association for the Advancement of Artificial Intelligence to illustrate the scale of the challenge. According to the survey, 90% spend more than 10% of their time on AI evaluation, 30% spend more than 30%, and 40% cite a lack of suitable evaluation methodologies as the biggest challenge.

These figures suggest testing remains labour-intensive even as AI adoption widens. They also highlight a persistent question for many organisations: how to establish a repeatable process for checking models, agents and other AI systems against business, safety and compliance expectations.

Black-box testing

FlightSim is based on black-box testing, which evaluates a system through its inputs and outputs rather than its internal structure. In practice, this means a company can test a model supplied by a vendor or partner without requiring disclosure of proprietary code.

This may be particularly relevant in commercial AI procurement, where vendors are often reluctant to expose internal workings but buyers still need evidence that a system performs as claimed. It also suits organisations using multiple models from different suppliers that want a common testing method.

The standalone launch allows companies to add FlightSim to existing deployment stacks and vendor management security processes. Rather than requiring customers to adopt the full governance platform, Monitaur is positioning the product as a focused evaluation layer that can be used independently.

That could broaden its reach beyond heavily regulated users already working with Monitaur. It may also appeal to developers who want an external validation process alongside their own testing, particularly for higher-risk use cases.

Monitaur's technology leadership said manual review and general-purpose testing become harder to scale as AI systems grow more complex.

"Objective validations that provide confidence and assurance in the trust and reliability of AI systems across different data and modeling modalities are very difficult to scale," said Dr. Andrew Clark, Co-Founder and Chief Technology Officer, Monitaur.

"As the proliferation and impacts of AI and agents continue to accelerate, human review capacity and generalized evaluation, testing, and validation techniques fall short. FlightSim is designed to address these gaps, and by making FlightSim a standalone offering, we enable a broader community to leverage our expertise, independent of our full platform," Clark said.

Wider backdrop

The product enters the market amid rising debate over safeguards for increasingly capable AI systems. Senior figures in the AI sector, including Sam Altman and Dario Amodei, have publicly argued that stronger oversight is needed as advanced models take on more significant tasks.

Against that backdrop, software suppliers are trying to build tools that support governance, auditability and procurement checks without forcing customers into one model provider or technical architecture. Monitaur's emphasis on external testing appears aimed at that cross-vendor environment.

FlightSim draws on research and established evaluation practices to assess AI systems before use. When used with the wider Monitaur platform, it can also feed governance automation and ongoing production monitoring, though the standalone version is now being offered separately.

Clark described the product as a response to the unpredictability of advanced systems and the need to assess them in real-world conditions beyond their original design assumptions.

"We have always been inspired by the foundations of aerospace engineering, complex system design, and concepts like Nassim Taleb's antifragility," Dr. Clark said. "Especially for our most consequential and high-impact systems, we should assume AI will encounter situations not anticipated, designed, or evaluated for. FlightSim complements developers' efforts and prepares our companies with a broader confidence and understanding of AI's fit for purpose and readiness before use."