trackmcp
Back to directory
XenoCoreGiger31

GEMMA-by-GOOGLE

View on GitHub

A fully local, autonomous AI penetration-testing agent powered by Gemma 4-12B and a Flask MCP tool server. Runs recon, attack loops, and report generation on its own — no cloud, no API keys. For authorized security testing only.

8 stars PythonOthers Updated Aug 18, 2026
autonomous-ai-agentscybersecurityethical-hackingexploitationgemmakali-linuxllmspenetration-testingpythonreconredteamsecurity-researchvulnerability-scanner

Documentation

https://github.com/user-attachments/assets/ba467fae-a4c9-4f63-b2e6-3fc30fb023f3


HALO is an autonomous security agent that runs inside a Linux environment driven

by a local LLM — Gemma 4-12B (uncensored / abliterated) served through

LM Studio. It plans, runs reconnaissance, chains attacks based on what it finds,

and writes a professional pentest report on its own. Everything runs locally:

no cloud, no API keys, nothing leaves your machine.

One word starts an engagement: `engage`.


What It Does

  • 🔍 Autonomous recon — masscan + nmap to discover open ports and services
  • ⚔️ Autonomous attack loop — selects and chains tools based on what it finds
  • 🌐 Web recon → attack pipeline — apex-to-URL enumeration (subdomains,

hosts, historical URLs), content discovery, template scanning and XSS, with

automatic flag capture on CTF-style web targets

  • Verified breaches, not banners — every attempt carries a single-use

challenge/nonce the exploit must echo *from inside the popped shell*; a bare

`uid=0` banner or a tarpit can't forge it, so a confirmed breach is a real one

(execution-derived evidence, consume-once at the gate)

  • 🎯 Curated PoC library — deterministic, self-evident exploits

(vsftpd 2.3.4, ingreslock, UnrealIRCd) fired through a sandboxed delivery

primitive that returns a real shell, not a guess

  • 🧠 Persistent negative-experience cache — learns what fails across *all*

sessions and stops wasting cycles on proven dead ends

  • 🧩 Adaptive skill injection — loads relevant attack playbooks into the

prompt based on the current goal

  • 📝 Automatic HTML reports — compiles findings into a branded report on exit
  • 🔒 100% local — Gemma 4-12B in LM Studio; nothing leaves your machine

Tool Arsenal

42 tools sit behind the agent's decision loop, all routed through the same

failure-caching layer. They are defined once in the `TOOLS` schema registry in

`halo_tools.py` and served over both transports (MCP and HTTP).

Recon & OSINT

ToolPurpose
`run_subfinder`Subdomain enumeration
`run_theharvester`Passive OSINT — emails, subdomains, hosts
`run_httpx`HTTP probing and fingerprinting
`run_katana`Web crawling
`run_sherlock`Username OSINT across 90+ platforms
`run_shodan`Internet-exposure intelligence lookups
`run_phoneinfoga`Phone-number OSINT
`run_phonextract`Phone-number OSINT / extraction
`run_ghosttrack`OSINT for username / IP / phone
`run_cloudfox`Cloud-infrastructure enumeration
`run_wafw00f`WAF / security-solution fingerprinting
`run_amass`Subdomain enumeration (passive by default)
`run_dnsx`DNS resolution and probing
`run_gau`Known URLs from OTX / Wayback / Common Crawl
`run_waybackurls`Historical URLs from the Wayback Machine
`run_gowitness`Web screenshotting for visual recon
`run_spiderfoot`Headless multi-module OSINT scanning
`run_recon_ng`recon-ng OSINT framework (non-interactive)

Scanning

ToolPurpose
`run_masscan`Fast port discovery
`run_nmap`Deep service/version scanning
`run_nikto`Web vulnerability scanning
`run_nuclei`Template-based vulnerability scanning
`run_netstat`Network connection analysis

Web & Fuzzing

ToolPurpose
`run_gobuster`Web directory brute forcing
`run_ffuf`Web fuzzing
`run_feroxbuster`Recursive content discovery
`run_dalfox`XSS scanning (reflected / stored / DOM)
`run_curl`HTTP request testing
`run_wget`File retrieval

Exploitation

ToolPurpose
`run_sqlmap`SQL injection testing
`run_searchsploit`Exploit lookup
`run_metasploit`Fire a chosen Metasploit module at a target (human-approved)
`run_exploit`Sandboxed execution of custom PoC scripts
`run_setoolkit`Social-engineering toolkit

Credentials

ToolPurpose
`run_hydra`Credential brute forcing
`run_ncrack`Network authentication cracking
`run_medusa`Fast parallel brute forcing
`run_john`Hash cracking

Enumeration & System

ToolPurpose
`run_enum4linux`SMB / Samba enumeration
`run_command`Arbitrary command execution
`read_file`Read file contents
`write_file`Write output to files

Architecture

A single tool engine (`halo_tools.py`) owns the arsenal and its schemas; two

thin transports sit on top of it, so the tools are defined exactly once:

code
agent_loop.py ──HTTP─►  tool_server.py ─┐
                                            ├─►  halo_tools.py  ──►  security tools
   MCP clients  ──stdio►  mcp_server.py  ──┘   (42-tool engine +
                                                 schema registry)
     │
     ├─►  agent_cache.py         (persistent negative-experience cache)
     ├─►  skills.py              (adaptive playbook injection)
     └─►  report_generator.py    (auto HTML pentest report on exit)
  • `mcp_server.py` — a spec-compliant Model Context Protocol server

(stdio, JSON-RPC 2.0). Point any MCP client (Claude Desktop, IDE agents,

inspectors) or an MCP registry at it to use HALO's arsenal as standard tools.

  • `tool_server.py` — the local Flask HTTP tool server (port 8000) the

autonomous agent loop drives.

Use HALO as an MCP server

jsonc
// e.g. an MCP client config
{
  "mcpServers": {
    "halo": { "command": "python3", "args": ["/abs/path/to/mcp_server.py"] }
  }
}

A ready-to-submit registry manifest lives in `server.json`.

Multi-agent layer

Engagements are coordinated by a set of specialist agents that pass a shared

message schema (`agent_schema.py`):

AgentRole
`planner_agent.py`Turns a goal into an ordered plan
`orchestrator_agent.py`Routes tasks to the right specialist
`vuln_discovery_agent.py`Surfaces candidate vulnerabilities
`attacker_agent.py`Branches into vuln-class specialists (SQLi, brute force, IDOR, SSRF, XSS, auth)
`validator_agent.py`Confirms findings against real evidence before they count
`debugger_agent.py`Diagnoses failed tool runs and adjusts

Sovereign Agent Layer

The negative-experience cache fingerprints every tool call. A call that fails

gets one retry; fail twice and it is blacklisted, so the agent moves on to a

more practical tool for the job. Over an engagement the agent structures its own

trial-and-error learning — building context, avoiding repeated dead ends, and

escalating intelligently — rather than re-running what it has already proven

doesn't work.

Verified breaches, not vibes

The hard problem with an autonomous attacker is knowing whether it *actually*

broke in or just parroted a hopeful banner. HALO answers this with a

challenge-response gate:

  • The orchestrator mints a per-attempt nonce, bound to that target and the

exact payload hash, before firing.

  • A breach only counts if the tool output carries a structured

`HALO-EVIDENCE nonce=… level=…` line echoing that nonce — which the

delivery primitive (`pocs/_delivery.py`) can only produce by running code

*inside* the shell it claims to have.

  • The nonce is consume-once: the gate (`exploitation_core.py:breach_confirmed`)

rejects a replayed or never-minted nonce, so a tarpit, a reflected string, or a

static `uid=0` banner cannot forge a confirmation.

The curated PoCs in `pocs/` are deterministic, self-evident bugs

(vsftpd 2.3.4, ingreslock 1524, UnrealIRCd 3.2.8.1) that pass this gate honestly —

they land a real root shell or they report nothing.


How It Was Built

HALO was built solo, from the ground up, in under six months by a self-taught

developer and security researcher. The multi-agent core came together one

specialist at a time, each verified against a real target before moving on:

  • Shared language: a common message schema (`agent_schema.py`) so the agents can talk to each other
  • Planner: turns a goal into an ordered plan, verified against live LM Studio
  • Orchestrator: routes each task to the right specialist
  • Vuln Discovery: surfaces candidate vulnerabilities, tested against a live Metasploitable target
  • Attacker: branches into SQLi / brute-force / IDOR / SSRF / XSS / auth specialists
  • Debugger: diagnoses failed tool runs and adjusts
  • Validator + reporting: findings are confirmed against real evidence before they count, then compiled into a client-readable report

From there the arsenal grew to 42 tools, a full web recon → attack pipeline with

flag capture, and challenge-response breach confirmation, while the

negative-experience cache turned trial-and-error into persistent learning across

sessions. Active development continues — new capabilities are pushed regularly;

see the changelog for the shipped milestones.


Stack

  • Model: Gemma 4-12B Instruct Abliterated (GGUF via LM Studio) — works with

any local model of your choosing

  • Agent: Python autonomous loop with MCP tool calls
  • Tool transports: a Model Context Protocol server (stdio) for MCP clients,

plus a Flask HTTP tool server on port 8000 for the agent loop

  • OS: Kali Linux (tested under UTM on Apple Silicon M1)
  • Hardware reference: MacBook Pro M1, 16 GB RAM

Quickstart

See **docs/QUICKSTART.md** for full setup. In short:

bash
git clone https://github.com/XenoCoreGiger31/GEMMA-by-GOOGLE.git
cd GEMMA-by-GOOGLE
python3 -m pip install -r requirements.txt

cp engagement.example.yaml engagement.yaml   # then fill in authorization + scope_targets

python3 tool_server.py      # terminal 1 — HTTP tool server on port 8000
python3 agent_loop.py       # terminal 2 — the agent

>>> engage 203.0.113.3     # full autonomous recon + attack
>>> run nmap on 10.0.0.1    # single-goal query
>>> exit                    # triggers HTML report generation

> Note: endpoints and paths default to a standard local setup (LM Studio on

> `localhost:1234`, HTTP tool server on `localhost:8000`). Override any of them

> with the `HALO_*` environment variables — see the

> environment overrides table. A few

> author-specific log/cache path defaults remain in `agent_cache.py` and

> `tool_server.py`; the env vars cover those too.

>

> `agent_loop.py` will not start without `engagement.yaml` — it's the

> authorization + scope gate every tool call passes through, not optional

> config. See step 5 of the Quickstart.


Running Tests

The unit tests use Python's built-in `unittest` — no extra dependencies:

bash
python3 -m unittest

Contributing

Contributions from the security, AI, and Python communities are welcome — see

CONTRIBUTING.md. Star the repo if it's useful to you, or open

a PR and let's build something together.

Actively developed by an independent, self-taught developer and security

researcher. New capabilities are pushed regularly.


This is a community project by an independent developer. It is **not affiliated

with, endorsed by, or sponsored by Google LLC.** "Gemma" is a trademark of

Google LLC.

> ⚠️ Content warning: The referenced model is heavily abliterated and will

> respond to sensitive requests without the usual guardrails. Use responsibly,

> in appropriate environments only.

> 🔒 Legal warning: This tool is intended strictly for authorized

> penetration testing and security research on systems you own or have

> explicit written permission to test. Unauthorized use is illegal.

License

Released under the MIT License.

Frequently asked questions

What is GEMMA-by-GOOGLE?

GEMMA-by-GOOGLE is A fully local, autonomous AI penetration-testing agent powered by Gemma 4-12B and a Flask MCP tool server. Runs recon, attack loops, and report generation on its own — no cloud, no API keys. For authorized security testing only.

How do I install GEMMA-by-GOOGLE?

Open the GitHub repository and follow its README. Most MCP servers are added to your client's MCP config, then called by your agent.

Is GEMMA-by-GOOGLE open source?

Yes — it is hosted on GitHub at https://github.com/XenoCoreGiger31/GEMMA-by-GOOGLE and has 8 stars.

Related MCP tools

Run your own MCP server? See who uses it and what to fix.

Measure it with TrackMCP