Proposal for feedback: CAV-Bench adapter for workflow commit validity #7159
Nixalkumar (Harimay23)
started this conversation in
Ideas
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
I am exploring an external Microsoft Agent Framework adapter for CAV-Bench, a deterministic benchmark that evaluates whether an agent action remained valid when it committed.
The benchmark separates expected outcome success from five execution-validity dimensions:
The initial Agent Framework scenario pack would focus on:
I am considering two implementation approaches:
Would one of these approaches align better with Agent Framework’s samples, AF Labs, or benchmarking work? I would also appreciate guidance on the most reliable workflow evidence for distinguishing attempted, committed, reconciled, and compensated operations.
This is an independent open-source project. I am requesting technical feedback, not endorsement or official integration.
Repository: Harimay23/cav-bench
Version DOI: 10.5281/zenodo.21364386
All reactions