GEMMA-by-GOOGLE
A fully local, autonomous AI penetration-testing agent powered by Gemma 4-12B and a Flask MCP tool server. Runs recon, attack loops, and report generation on its own — no cloud, no API keys. For authorized security testing only.
Documentation
https://github.com/user-attachments/assets/ba467fae-a4c9-4f63-b2e6-3fc30fb023f3
HALO is an autonomous security agent that runs inside a Linux environment driven
by a local LLM — Gemma 4-12B (uncensored / abliterated) served through
LM Studio. It plans, runs reconnaissance, chains attacks based on what it finds,
and writes a professional pentest report on its own. Everything runs locally:
no cloud, no API keys, nothing leaves your machine.
One word starts an engagement: `engage`.
What It Does
- 🔍 Autonomous recon — masscan + nmap to discover open ports and services
- ⚔️ Autonomous attack loop — selects and chains tools based on what it finds
- 🌐 Web recon → attack pipeline — apex-to-URL enumeration (subdomains,
hosts, historical URLs), content discovery, template scanning and XSS, with
automatic flag capture on CTF-style web targets
- ✅ Verified breaches, not banners — every attempt carries a single-use
challenge/nonce the exploit must echo *from inside the popped shell*; a bare
`uid=0` banner or a tarpit can't forge it, so a confirmed breach is a real one
(execution-derived evidence, consume-once at the gate)
- 🎯 Curated PoC library — deterministic, self-evident exploits
(vsftpd 2.3.4, ingreslock, UnrealIRCd) fired through a sandboxed delivery
primitive that returns a real shell, not a guess
- 🧠 Persistent negative-experience cache — learns what fails across *all*
sessions and stops wasting cycles on proven dead ends
- 🧩 Adaptive skill injection — loads relevant attack playbooks into the
prompt based on the current goal
- 📝 Automatic HTML reports — compiles findings into a branded report on exit
- 🔒 100% local — Gemma 4-12B in LM Studio; nothing leaves your machine
Tool Arsenal
42 tools sit behind the agent's decision loop, all routed through the same
failure-caching layer. They are defined once in the `TOOLS` schema registry in
`halo_tools.py` and served over both transports (MCP and HTTP).
Recon & OSINT
| Tool | Purpose |
|---|---|
| `run_subfinder` | Subdomain enumeration |
| `run_theharvester` | Passive OSINT — emails, subdomains, hosts |
| `run_httpx` | HTTP probing and fingerprinting |
| `run_katana` | Web crawling |
| `run_sherlock` | Username OSINT across 90+ platforms |
| `run_shodan` | Internet-exposure intelligence lookups |
| `run_phoneinfoga` | Phone-number OSINT |
| `run_phonextract` | Phone-number OSINT / extraction |
| `run_ghosttrack` | OSINT for username / IP / phone |
| `run_cloudfox` | Cloud-infrastructure enumeration |
| `run_wafw00f` | WAF / security-solution fingerprinting |
| `run_amass` | Subdomain enumeration (passive by default) |
| `run_dnsx` | DNS resolution and probing |
| `run_gau` | Known URLs from OTX / Wayback / Common Crawl |
| `run_waybackurls` | Historical URLs from the Wayback Machine |
| `run_gowitness` | Web screenshotting for visual recon |
| `run_spiderfoot` | Headless multi-module OSINT scanning |
| `run_recon_ng` | recon-ng OSINT framework (non-interactive) |
Scanning
| Tool | Purpose |
|---|---|
| `run_masscan` | Fast port discovery |
| `run_nmap` | Deep service/version scanning |
| `run_nikto` | Web vulnerability scanning |
| `run_nuclei` | Template-based vulnerability scanning |
| `run_netstat` | Network connection analysis |
Web & Fuzzing
| Tool | Purpose |
|---|---|
| `run_gobuster` | Web directory brute forcing |
| `run_ffuf` | Web fuzzing |
| `run_feroxbuster` | Recursive content discovery |
| `run_dalfox` | XSS scanning (reflected / stored / DOM) |
| `run_curl` | HTTP request testing |
| `run_wget` | File retrieval |
Exploitation
| Tool | Purpose |
|---|---|
| `run_sqlmap` | SQL injection testing |
| `run_searchsploit` | Exploit lookup |
| `run_metasploit` | Fire a chosen Metasploit module at a target (human-approved) |
| `run_exploit` | Sandboxed execution of custom PoC scripts |
| `run_setoolkit` | Social-engineering toolkit |
Credentials
| Tool | Purpose |
|---|---|
| `run_hydra` | Credential brute forcing |
| `run_ncrack` | Network authentication cracking |
| `run_medusa` | Fast parallel brute forcing |
| `run_john` | Hash cracking |
Enumeration & System
| Tool | Purpose |
|---|---|
| `run_enum4linux` | SMB / Samba enumeration |
| `run_command` | Arbitrary command execution |
| `read_file` | Read file contents |
| `write_file` | Write output to files |
Architecture
A single tool engine (`halo_tools.py`) owns the arsenal and its schemas; two
thin transports sit on top of it, so the tools are defined exactly once:
agent_loop.py ──HTTP─► tool_server.py ─┐
├─► halo_tools.py ──► security tools
MCP clients ──stdio► mcp_server.py ──┘ (42-tool engine +
schema registry)
│
├─► agent_cache.py (persistent negative-experience cache)
├─► skills.py (adaptive playbook injection)
└─► report_generator.py (auto HTML pentest report on exit)- `mcp_server.py` — a spec-compliant Model Context Protocol server
(stdio, JSON-RPC 2.0). Point any MCP client (Claude Desktop, IDE agents,
inspectors) or an MCP registry at it to use HALO's arsenal as standard tools.
- `tool_server.py` — the local Flask HTTP tool server (port 8000) the
autonomous agent loop drives.
Use HALO as an MCP server
// e.g. an MCP client config
{
"mcpServers": {
"halo": { "command": "python3", "args": ["/abs/path/to/mcp_server.py"] }
}
}A ready-to-submit registry manifest lives in `server.json`.
Multi-agent layer
Engagements are coordinated by a set of specialist agents that pass a shared
message schema (`agent_schema.py`):
| Agent | Role |
|---|---|
| `planner_agent.py` | Turns a goal into an ordered plan |
| `orchestrator_agent.py` | Routes tasks to the right specialist |
| `vuln_discovery_agent.py` | Surfaces candidate vulnerabilities |
| `attacker_agent.py` | Branches into vuln-class specialists (SQLi, brute force, IDOR, SSRF, XSS, auth) |
| `validator_agent.py` | Confirms findings against real evidence before they count |
| `debugger_agent.py` | Diagnoses failed tool runs and adjusts |
Sovereign Agent Layer
The negative-experience cache fingerprints every tool call. A call that fails
gets one retry; fail twice and it is blacklisted, so the agent moves on to a
more practical tool for the job. Over an engagement the agent structures its own
trial-and-error learning — building context, avoiding repeated dead ends, and
escalating intelligently — rather than re-running what it has already proven
doesn't work.
Verified breaches, not vibes
The hard problem with an autonomous attacker is knowing whether it *actually*
broke in or just parroted a hopeful banner. HALO answers this with a
challenge-response gate:
- The orchestrator mints a per-attempt nonce, bound to that target and the
exact payload hash, before firing.
- A breach only counts if the tool output carries a structured
`HALO-EVIDENCE nonce=… level=…` line echoing that nonce — which the
delivery primitive (`pocs/_delivery.py`) can only produce by running code
*inside* the shell it claims to have.
- The nonce is consume-once: the gate (`exploitation_core.py:breach_confirmed`)
rejects a replayed or never-minted nonce, so a tarpit, a reflected string, or a
static `uid=0` banner cannot forge a confirmation.
The curated PoCs in `pocs/` are deterministic, self-evident bugs
(vsftpd 2.3.4, ingreslock 1524, UnrealIRCd 3.2.8.1) that pass this gate honestly —
they land a real root shell or they report nothing.
How It Was Built
HALO was built solo, from the ground up, in under six months by a self-taught
developer and security researcher. The multi-agent core came together one
specialist at a time, each verified against a real target before moving on:
- Shared language: a common message schema (`agent_schema.py`) so the agents can talk to each other
- Planner: turns a goal into an ordered plan, verified against live LM Studio
- Orchestrator: routes each task to the right specialist
- Vuln Discovery: surfaces candidate vulnerabilities, tested against a live Metasploitable target
- Attacker: branches into SQLi / brute-force / IDOR / SSRF / XSS / auth specialists
- Debugger: diagnoses failed tool runs and adjusts
- Validator + reporting: findings are confirmed against real evidence before they count, then compiled into a client-readable report
From there the arsenal grew to 42 tools, a full web recon → attack pipeline with
flag capture, and challenge-response breach confirmation, while the
negative-experience cache turned trial-and-error into persistent learning across
sessions. Active development continues — new capabilities are pushed regularly;
see the changelog for the shipped milestones.
Stack
- Model: Gemma 4-12B Instruct Abliterated (GGUF via LM Studio) — works with
any local model of your choosing
- Agent: Python autonomous loop with MCP tool calls
- Tool transports: a Model Context Protocol server (stdio) for MCP clients,
plus a Flask HTTP tool server on port 8000 for the agent loop
- OS: Kali Linux (tested under UTM on Apple Silicon M1)
- Hardware reference: MacBook Pro M1, 16 GB RAM
Quickstart
See **docs/QUICKSTART.md** for full setup. In short:
git clone https://github.com/XenoCoreGiger31/GEMMA-by-GOOGLE.git
cd GEMMA-by-GOOGLE
python3 -m pip install -r requirements.txt
cp engagement.example.yaml engagement.yaml # then fill in authorization + scope_targets
python3 tool_server.py # terminal 1 — HTTP tool server on port 8000
python3 agent_loop.py # terminal 2 — the agent
>>> engage 203.0.113.3 # full autonomous recon + attack
>>> run nmap on 10.0.0.1 # single-goal query
>>> exit # triggers HTML report generation> Note: endpoints and paths default to a standard local setup (LM Studio on
> `localhost:1234`, HTTP tool server on `localhost:8000`). Override any of them
> with the `HALO_*` environment variables — see the
> environment overrides table. A few
> author-specific log/cache path defaults remain in `agent_cache.py` and
> `tool_server.py`; the env vars cover those too.
>
> `agent_loop.py` will not start without `engagement.yaml` — it's the
> authorization + scope gate every tool call passes through, not optional
> config. See step 5 of the Quickstart.
Running Tests
The unit tests use Python's built-in `unittest` — no extra dependencies:
python3 -m unittestContributing
Contributions from the security, AI, and Python communities are welcome — see
CONTRIBUTING.md. Star the repo if it's useful to you, or open
a PR and let's build something together.
Actively developed by an independent, self-taught developer and security
researcher. New capabilities are pushed regularly.
Disclaimer & Legal
This is a community project by an independent developer. It is **not affiliated
with, endorsed by, or sponsored by Google LLC.** "Gemma" is a trademark of
Google LLC.
> ⚠️ Content warning: The referenced model is heavily abliterated and will
> respond to sensitive requests without the usual guardrails. Use responsibly,
> in appropriate environments only.
> 🔒 Legal warning: This tool is intended strictly for authorized
> penetration testing and security research on systems you own or have
> explicit written permission to test. Unauthorized use is illegal.
License
Released under the MIT License.
Frequently asked questions
What is GEMMA-by-GOOGLE?
GEMMA-by-GOOGLE is A fully local, autonomous AI penetration-testing agent powered by Gemma 4-12B and a Flask MCP tool server. Runs recon, attack loops, and report generation on its own — no cloud, no API keys. For authorized security testing only.
How do I install GEMMA-by-GOOGLE?
Open the GitHub repository and follow its README. Most MCP servers are added to your client's MCP config, then called by your agent.
Is GEMMA-by-GOOGLE open source?
Yes — it is hosted on GitHub at https://github.com/XenoCoreGiger31/GEMMA-by-GOOGLE and has 8 stars.
Related MCP tools
AI-powered OSINT agent with interactive REPL, MCP server, and CLI. 19 tools. Works with Claude, GPT-4, or local models. For authorized security research only.
Production-grade MCP server giving Claude 27 security intelligence tools across 21 APIs — CVE lookup, EPSS scoring, CISA KEV, MITRE ATT&CK, Shodan, VirusTotal, and more.
An LLM agent that conducts deep research (local and web) on any given topic and generates a long report with citations. Built for the Model Context Protocol to
Automate browser based workflows with AI
🚀 The fast, Pythonic way to build MCP servers and clients Trusted by 19900+ developers. Trusted by 19900+ developers. Trusted by 19900+ developers.
Python SQL Parser and Transpiler
Run your own MCP server? See who uses it and what to fix.
Measure it with TrackMCP