The model is not the agent.
Backbone or Backbone-Architecture? A Controlled Study of LLM Agents on Web Penetration-Testing CTFs
A controlled study, by Cyphlon CEO Zeeshan Sultan, of what drives LLM agent performance in web penetration testing: the model, or the scaffold around it. Across the 104 single-flag CTFs of the XBOW validation suite, 9 agent scaffolds and 10 backbone models, more than 2,700 runs were scored on exact flag retrieval and every solve was audited. Holding the model fixed and changing only the scaffold moved the same backbone from 0 to 49 solves.
104 web CTFs
The XBOW validation suite of single-flag web-exploitation challenges.
9 scaffolds, 10 models
Agent scaffolds and backbone models varied independently, to separate their effects.
2,700+ audited runs
Every transcript captured and scored on exact flag retrieval, with 142 solves audited.
The scaffold decides
Same model, different scaffold: from 0 to 49 solves on the same challenges.
More architecture can hurt
In a controlled ablation, a bespoke planner scored 72 against 90 for a leaner agent on the same model.
Open and reproducible
Code, data and the paper are public on GitHub.