WAO logoWaltrump AI Orchestrator

Application and agent evaluation

Evaluate what an agent proposes, changes and hands off—not only its final answer.

Review versioned application and agent scenarios, task outcomes, tool proposals, recovery, handoff, safety and human approval without automatic tool execution.

What the packs examine

Representative scenarios with a controlled action boundary.

WAO provides ten versioned starter pack definitions with normal, adversarial and failure scenarios. Installation prepares evaluation structure; it does not activate a provider or run a tool.

01

Task completion

Did the workflow reach the defined outcome inside the scenario boundary?

02

Tool proposal

Was the proposed tool and parameter shape appropriate, allowed and reviewable?

03

Retries and recovery

Did the workflow respond safely to failure, timeout or unavailable dependencies?

04

Human intervention

Was approval requested at the right point, with enough information for an accountable decision?

05

Handoff quality

Did the workflow preserve context, limitations and next action when responsibility changed?

06

Evidence coverage

Can each conclusion be linked to an approved scenario, bounded trace and customer-safe result?

Fail-closed agent boundary

A proposal is not an action.

The currently reviewed live path is one exact text-to-text proposal-only tuple with an accepted limited live case and deterministic security coverage.

WAO never executes a proposed tool automatically. Bounded human approval, an isolated sandbox, an approved tool schema, exact activation and project BYOK are required.
Exact provider/model/operation tupleCustomer-managed project credentialBounded requests, spend and concurrencyNo hidden reasoning retentionNo raw tool arguments in public evidenceNo automatic provider promotion

Acceptance truth

Review access exists; completion evidence remains scoped.

Application and Agent Evaluation Packs remain REVIEW_AVAILABLE. Three owner-operated end-to-end pack runs are still recorded as an acceptance gate, so this page does not claim completed customer value or production readiness.

4 capability records

REVIEW_AVAILABLEREQUIRES_ACTIVATIONREQUIRES_BYOK
Text, metadata or approved evidence

AI Application and Agent Evaluation Packs

Can this application or proposal-only agent complete representative scenarios safely?

Evidence
Versioned scenario, bounded trace and human-review evidence
Credential boundary
Customer-managed provider credential encrypted in the exact project is required only for approved provider execution
Provider/model boundary
Provider-neutral; exact combinations are separately qualified · Workload-selected and entitlement controlled
Execution dependency
exact project byok and tuple activation for provider or agent runs
Known limitation
Ten governed pack definitions are reviewable. Three owner-operated end-to-end pack runs remain an acceptance gate; installation alone executes nothing.
Request this path
CONTROLLED_BETA
Text, metadata or approved evidence

AI Safety

What safety signals and unresolved risks need review?

Evidence
Capability-specific evidence and known limitations
Credential boundary
No provider credential is needed unless a separately approved workflow executes a provider
Provider/model boundary
Provider-neutral; exact combinations are separately qualified · Workload-selected and entitlement controlled
Execution dependency
none
Known limitation
Safety checks are evidence and control inputs, not a guarantee that an AI system is safe.
Request this path
REQUIRES_ACTIVATIONREQUIRES_BYOKREVIEW_AVAILABLE
Text and approved tool schema

Agent and Tool Intelligence

Does the agent complete the task safely with auditable tool use?

Evidence
Capability-specific evidence and known limitations
Credential boundary
Customer-managed provider credential encrypted in the exact project is required only for approved provider execution
Provider/model boundary
One exact Google Gemini proposal-only agent tuple · Exact model/version confirmed during controlled access review
Execution dependency
exact project byok and tuple activation proposal only
Known limitation
One exact Google Gemini text-to-text agent.run proposal path has controlled-beta certification with an accepted limited live case. WAO never executes a proposed tool automatically; bounded human approval remains required.
Request this path
CONTROLLED_BETA
Text, metadata or approved evidence

Human Review

Where is accountable human judgment required?

Evidence
Independent review and adjudication record
Credential boundary
No provider credential is needed unless a separately approved workflow executes a provider
Provider/model boundary
Provider-neutral; exact combinations are separately qualified · Workload-selected and entitlement controlled
Execution dependency
none
Known limitation
Human review supports a decision; it does not replace representative data, safety controls or accountable approval.
Request this path

Controlled Beta Access

Define the task, allowed tools and human decision before execution.

Bring synthetic or customer-safe scenarios. Provider execution remains off until the exact reviewed tuple is activated.