Train
Prepare validated datasets and investigate supervised fine-tuning with efficient LoRA / QLoRA methods.
Researching, training, and evaluating intelligent agents in controlled environments. Efficient models. Reproducible experiments. A more accessible path to security automation.
Specialized security tasks demand reliable reasoning, tool use, and measurable outcomes. Developing efficient agents requires more than connecting a model to a set of tools.
From carefully prepared examples to independently verified outcomes: each stage informs the next.
Prepare validated datasets and investigate supervised fine-tuning with efficient LoRA / QLoRA methods.
Measure task completion, steps, invalid actions, and policy rejections in reproducible synthetic environments.
Inspect failures and execution trajectories to refine data, training, and agent behavior.
Investigate efficient inference architectures for accessible infrastructure. A research direction, not a released product.
An exploratory comparison of two recorded Qwen runs on the same synthetic task seeds. Report labels are used as recorded; they are not independently certified model identities.
| Recorded run | Successes | Mean steps | Invalid actions | Policy rejections |
|---|---|---|---|---|
| qwen35-base-100-199 | 41 / 100 | 11.56 | 0 | 76 |
| qwen35-sft-v2-val | 63 / 100 | 11.35 | 83 | 0 |
Both reports use seeds 100–199, split “train”, benchmark 0.3.1, a 180-second provider timeout, a 256-token completion budget, and step limits of 10 / 15 / 20 by difficulty. There are no recorded provider or infrastructure errors. Mean steps include successes and failures. Policy rejections count attempted actions blocked by the laboratory policy; they do not imply successful boundary violations.
These are single runs on 100 synthetic tasks, within the benchmark training seed range. Training overlap and full inference equivalence have not been independently audited. The 22 percentage point difference is descriptive, not a causal SFT improvement claim or evidence of unseen-task generalization. No production security performance or cost savings are established.
Download sanitized evidence ↓Only allowlisted aggregate data is included. Raw trajectories, internal configuration, and private datasets are excluded.
The repository contains a policy-enforced synthetic task environment, an independent outcome verifier, a sequential teacher collection pipeline, validated offline dataset preparation, and an optional SFT / QLoRA training harness. These are research components, not a finished commercial platform.
A conceptual architecture for tool-using agents. Restricted tools and independent verification separate agent decisions from evaluation outcomes.
Potential future applications. These are research goals, not available commercial services.
Explore agents that identify and correct configuration issues within explicitly authorized scopes.
Investigate AI-assisted testing in controlled, synthetic environments.
Study multi-step reasoning and tool use through inspectable agent trajectories.
Develop reproducible evaluations with transparent conditions and independent verification.
Investigate smaller models as a practical foundation for domain-specific automation.
A future where cybersecurity agents can be developed and evaluated without depending exclusively on enormous models and prohibitively expensive infrastructure.
Cyber-Agent-Lab is an early-stage, self-funded research and engineering project exploring efficient autonomous agents for cybersecurity. Its focus is controlled experimentation, reproducible evaluation, and a responsible path toward useful security automation.
For future research collaboration and technical conversations.