Small, Free, and Effective: Orchestrating Open-Weight Small Language Models to Outperform Single LLM for Malware Analysis
Merged summary
TL;DR - Orchestrating open-weight small language models substantially improves malware-report analysis, with a hybrid evidence-and-debate system narrowly outperforming the best ungrounded frontier LLM baseline.
- Evaluated 11 open-weight SLMs, three cybersecurity-pretrained models, and six frontier LLMs on Meta’s CyberSecEval Malware Analysis benchmark.
- Tested multi-agent pipelines, adversarial debate, hierarchical expert consultation, and a hybrid orchestration architecture.
- The Qwen3-4B and Foundation-Sec-8B hybrid reached 35.30% accuracy, versus 22.54% for the strongest cybersecurity baseline and 34.77% for the strongest ungrounded frontier model.
- Grounded Gemini remained best at 38.22%, highlighting the value of structured evidence collection across model sizes.
Sources (1)
Small, Free, and Effective: Orchestrating Open-Weight Small Language Models to Outperform Single LLM for Malware Analysis
TL;DR - Orchestrating open-weight small language models substantially improves malware-report analysis, with a hybrid evidence-and-debate system narrowly outperforming the best ungrounded frontier LLM baseline.
- Evaluated 11 open-weight SLMs, three cybersecurity-pretrained models, and six frontier LLMs on Meta’s CyberSecEval Malware Analysis benchmark.
- Tested multi-agent pipelines, adversarial debate, hierarchical expert consultation, and a hybrid orchestration architecture.
- The Qwen3-4B and Foundation-Sec-8B hybrid reached 35.30% accuracy, versus 22.54% for the strongest cybersecurity baseline and 34.77% for the strongest ungrounded frontier model.
- Grounded Gemini remained best at 38.22%, highlighting the value of structured evidence collection across model sizes.