Independent research · Early stage

Building the next generation of autonomous cybersecurity agents.

Researching, training, and evaluating intelligent agents in controlled environments. Efficient models. Reproducible experiments. A more accessible path to security automation.

CONCEPTUAL AGENT LOOP: AI agent, Tool interface, Security environment, Evaluation, FeedbackTool interfaceEvaluationSecurity environmentFeedbackAI agent
CONCEPTUAL AGENT LOOP01 → 06
Efficient AI modelsValidated datasetsReproducible evaluationControlled environments
01 / The problem

Capability should not require unlimited compute.

Specialized security tasks demand reliable reasoning, tool use, and measurable outcomes. Developing efficient agents requires more than connecting a model to a set of tools.

  1. 01Large models can make experimentation costly.
  2. 02Agent evaluations are difficult to reproduce.
  3. 03Smaller models struggle with complex technical workflows.
  4. 04Security research needs safe, authorized environments.
02 / Our approach

A research loop, built to learn.

From carefully prepared examples to independently verified outcomes: each stage informs the next.

Train

Prepare validated datasets and investigate supervised fine-tuning with efficient LoRA / QLoRA methods.

Evaluate

Measure task completion, steps, invalid actions, and policy rejections in reproducible synthetic environments.

Improve

Inspect failures and execution trajectories to refine data, training, and agent behavior.

Deploy

Investigate efficient inference architectures for accessible infrastructure. A research direction, not a released product.

03 / Research & benchmarks

Measured progress.
Visible limitations.

An exploratory comparison of two recorded Qwen runs on the same synthetic task seeds. Report labels are used as recorded; they are not independently certified model identities.

Benchmark v0.3.1 · 2026-10-10 UTC
Task success rate
Qwen · Base41%
Qwen · SFT v263%
Task success rate · 100 episodes per run
Recorded runSuccessesMean stepsInvalid actionsPolicy rejections
qwen35-base-100-19941 / 10011.56076
qwen35-sft-v2-val63 / 10011.35830
Evaluation conditions & limitations

Both reports use seeds 100–199, split “train”, benchmark 0.3.1, a 180-second provider timeout, a 256-token completion budget, and step limits of 10 / 15 / 20 by difficulty. There are no recorded provider or infrastructure errors. Mean steps include successes and failures. Policy rejections count attempted actions blocked by the laboratory policy; they do not imply successful boundary violations.

These are single runs on 100 synthetic tasks, within the benchmark training seed range. Training overlap and full inference equivalence have not been independently audited. The 22 percentage point difference is descriptive, not a causal SFT improvement claim or evidence of unseen-task generalization. No production security performance or cost savings are established.

Download sanitized evidence ↓

Only allowlisted aggregate data is included. Raw trajectories, internal configuration, and private datasets are excluded.

What is implemented

The repository contains a policy-enforced synthetic task environment, an independent outcome verifier, a sequential teacher collection pipeline, validated offline dataset preparation, and an optional SFT / QLoRA training harness. These are research components, not a finished commercial platform.

04 / Technology

Autonomy within clear boundaries.

A conceptual architecture for tool-using agents. Restricted tools and independent verification separate agent decisions from evaluation outcomes.

01Language model
02AI agent
03Tool interface
04Security environment
05Evaluation
06Feedback
05 / Applications

Research today.
Possibilities ahead.

Potential future applications. These are research goals, not available commercial services.

01

Security configuration auditing

Explore agents that identify and correct configuration issues within explicitly authorized scopes.

02

Defensive security testing

Investigate AI-assisted testing in controlled, synthetic environments.

03

Autonomous security research

Study multi-step reasoning and tool use through inspectable agent trajectories.

04

Cybersecurity agent benchmarking

Develop reproducible evaluations with transparent conditions and independent verification.

05

Efficient specialized AI

Investigate smaller models as a practical foundation for domain-specific automation.

06 / Our vision

Specialized intelligence.
More accessible by design.

A future where cybersecurity agents can be developed and evaluated without depending exclusively on enormous models and prohibitively expensive infrastructure.

07 / About the lab

Independent research.
Intentional engineering.

Cyber-Agent-Lab is an early-stage, self-funded research and engineering project exploring efficient autonomous agents for cybersecurity. Its focus is controlled experimentation, reproducible evaluation, and a responsible path toward useful security automation.

Help shape what comes next.

For future research collaboration and technical conversations.

Get in touch · LinkedIn ↗