Sigma launches independent verification for enterprise AI agents

Sep. 21, 2026
By AI, Created 23:00 UTC, Sep 21, 2026, AGP -

Sigma has launched Sigma Eval, a human-supervised verification service for conversational AI agents that checks safety, user experience and performance before and after deployment. The move targets growing enterprise concerns that AI systems are going live faster than they can be independently tested.

Why it matters: - Enterprise AI agents are moving into customer service, sales, claims handling and internal support faster than many teams can independently verify them. - Sigma Eval is designed to give boards, regulators and customers evidence that an AI agent behaves safely in real-world use, not just in a controlled demo. - The service aims to reduce business risk tied to hallucinations, data exposure, broken user journeys and post-launch drift.

What happened: - Sigma announced Sigma Eval on September 22, 2026, in London. - The service provides independent assurance for conversational AI systems before and after deployment. - Sigma says the verification layer evaluates agents across 12 dimensions covering safety, user experience and performance. - Sigma is opening Sigma Eval with a free evaluation report for any organization operating a conversational agent. - The company is also offering a limited number of repeat evaluations to check whether remediation work improved results.

The details: - Sigma Eval measures the behavior of the agent, not the model builder, prompt author or system operator. - The service is meant to run continuously across the AI lifecycle. - Before deployment, it tests how an agent handles adversarial users, sensitive data, off-topic pressure and unresolved requests. - After deployment, it monitors for drift as prompts change, models are upgraded, knowledge bases shift and integrations move. - The 12 dimensions are grouped into three categories: safety, user experience and performance. - Safety covers bias, toxicity, hallucinations, opacity, PII exposure and vulnerability to manipulation. - User experience covers context misalignment, conversational inconsistency and user disengagement. - Performance covers trajectory performance, customer effort and resolution cost. - Sigma says its evidence comes from a proprietary system that combines synthetic users with synthetic conversations and can test agents at scale in any language. - Clients receive a scorecard with separate results for each dimension rather than one blended score. - The service requires no integration work, code access or model access. - Organizations submit a public link to the conversational agent plus basic contact details and context such as purpose, languages, KPIs and known edge cases. - Sigma Cognition runs the evaluation and sends a confidential report by email. - Reports are shared only with the requesting organization and are not used as public benchmarking material or promotional content. - Sigma says its verification work builds on applied research in AI safety, evaluation and anonymization with academic and institutional partners. - That research also underpins Sigma Cypher, which removes personally identifiable information before analysis. - Sigma Group says it develops AI products focused on data quality, multimodal anonymization, generative AI and trusted data platforms. - The company says its core offerings include Sigma Eval, Sigma Cypher and Sigma Verified.

Between the lines: - The launch reflects a broader governance shift: enterprises are being pushed to prove AI systems work as intended after deployment, not just during internal testing. - Sigma is positioning itself as an outside verifier rather than another AI vendor, which is meant to make the findings easier for third parties to trust. - The emphasis on human supervision suggests the market is not yet ready to rely on fully automated AI auditing alone. - The free report offer looks like a low-friction way to seed adoption while building demand for repeat checks. - Two survey snapshots cited by Sigma point to a gap between production use and evaluation practices: 57% of more than 1,300 professionals said their organizations have AI agents in production, while 52% reported offline evaluations; in a separate survey of 157 enterprises, 50% said they had deployed an AI agent or LLM feature that later caused a customer-facing failure after passing internal checks.

What’s next: - Sigma says it will continue offering repeat evaluations so organizations can measure the impact of remediation work over time. - The company is likely betting that recurring verification becomes a standard layer in enterprise AI governance as agents change after launch. - Sigma’s broader research base suggests additional verification and safety products may follow as demand for independent testing grows.

The bottom line: - Sigma is trying to make AI assurance look more like established oversight in finance, safety and security: independent, documented and repeatable.

Disclaimer: This article was produced by AGP Wire with the assistance of artificial intelligence based on original source content and has been refined to improve clarity, structure, and readability. This content is provided on an “as is” basis. While care has been taken in its preparation, it may contain inaccuracies or omissions, and readers should consult the original source and independently verify key information where appropriate. This content is for informational purposes only and does not constitute legal, financial, investment, or other professional advice.

Sign up for:

Applied Technology News

The daily local news briefing you can trust. Every day. Subscribe now.

By signing up, you agree to our Terms & Conditions.

Share this page:

Advanced Search Options

Search for:

Search scope:

Type:

Search in:

Date range:

The last

Sort by:

Sign up for:

Applied Technology News

The daily local news briefing you can trust. Every day. Subscribe now.

By signing up, you agree to our Terms & Conditions.