Dear CVE secretariat,
we are submitting the following session proposal:
Format: Lightning Talk
Theme Alignment: Representing uncertainty and abstraction more clearly
Contributors: Obinna Okeke, Sushant Nepal, Venkat Sai Suman Lamba Karanam, Yan Wu
Abstract:
We evaluated GPT-4o on the SecVulEval benchmark across 145 CWE categories, measuring the gap between binary vulnerability detection (vulnerable vs. safe) and exact CWE classification. We introduced a "Family Accuracy" metric and found that while exact CWE classification accuracy was low, models correctly identified the broader CWE family (parent/child/sibling relationships) far more often — showing that most classification errors are near misses between closely related CWEs rather than random or cross-family mistakes.
Possible discussion questions could include:
- How should CVE/CWE tooling represent classification confidence when a model identifies the correct family but not the exact CWE?
- What structural or hierarchical CWE information could be surfaced during CWE assignment to reduce near-miss errors?
- How does prompt-only LLM classification compare to fine-tuned or retrieval-augmented approaches for CWE assignment at scale?
Full details are available in our published paper, "From Binary Vulnerability Detection to CWE Classification: A Hierarchical Prompting Study" (ISDFS 2026, IEEE): https://ieeexplore.ieee.org/document/11459116
I know I'm submitting this after the July 14 deadline, but would be happy to be considered if there is still room on the agenda.
Dear CVE secretariat,
we are submitting the following session proposal:
Format: Lightning Talk
Theme Alignment: Representing uncertainty and abstraction more clearly
Contributors: Obinna Okeke, Sushant Nepal, Venkat Sai Suman Lamba Karanam, Yan Wu
Abstract:
We evaluated GPT-4o on the SecVulEval benchmark across 145 CWE categories, measuring the gap between binary vulnerability detection (vulnerable vs. safe) and exact CWE classification. We introduced a "Family Accuracy" metric and found that while exact CWE classification accuracy was low, models correctly identified the broader CWE family (parent/child/sibling relationships) far more often — showing that most classification errors are near misses between closely related CWEs rather than random or cross-family mistakes.
Possible discussion questions could include:
Full details are available in our published paper, "From Binary Vulnerability Detection to CWE Classification: A Hierarchical Prompting Study" (ISDFS 2026, IEEE): https://ieeexplore.ieee.org/document/11459116
I know I'm submitting this after the July 14 deadline, but would be happy to be considered if there is still room on the agenda.