trackmcp
Back to directory

Protocol for Augmented Memory of Project Artifacts (MCP compatible)

30 stars JavaScriptOthers Updated Sep 2, 2026
aiaugmenteddeveloperiamcpmemoryprojectprotocolsoftwaredevelopment

Documentation

PAMPA – Protocol for Augmented Memory of Project Artifacts

Version 1.12.x Β· Semantic Search Β· MCP Compatible Β· Node.js

Give your AI agents an always-updated, queryable memory of any codebase – with intelligent semantic search and automatic learning – in one `npx` command.

> πŸ‡ͺπŸ‡Έ **VersiΓ³n en EspaΓ±ol | πŸ‡ΊπŸ‡Έ English Version | πŸ€– Agent Version**

🌟 What's New in v1.12 - Advanced Search & Multi-Project Support

🎯 Scoped Search Filters - Filter by `path_glob`, `tags`, `lang` for precise results

πŸ”„ Hybrid Search - BM25 + Vector fusion with reciprocal rank blending (enabled by default)

🧠 Cross-Encoder Re-Ranker - Transformers.js reranker for precision boosts

πŸ‘€ File Watcher - Real-time incremental indexing with Merkle-like hashing

πŸ“¦ Context Packs - Reusable search scopes with CLI + MCP integration

πŸ› οΈ Multi-Project CLI - `--project` and `--directory` aliases for clarity

πŸ† **Performance Analysis** - Architectural comparison with general-purpose IDE tools

Major improvements:

  • 40% faster indexing with incremental updates
  • 60% better precision with hybrid search + reranker
  • 3x faster multi-project operations with explicit paths
  • 90% reduction in duplicate function creation with symbol boost
  • Specialized architecture for semantic code search

🌟 Why PAMPA?

Large language model agents can read thousands of tokens, but projects easily reach millions of characters. Without an intelligent retrieval layer, agents:

  • Recreate functions that already exist
  • Misname APIs (newUser vs. createUser)
  • Waste tokens loading repetitive code (`vendor/`, `node_modules/`...)
  • Fail when the repository grows

PAMPA solves this by turning your repository into a semantic code memory graph:

1. Chunking – Each function/class becomes an atomic chunk

2. Semantic Tagging – Automatic extraction of semantic tags from code context

3. Embedding – Enhanced chunks are vectorized with advanced embedding models

4. Learning – System learns from successful searches and caches intentions

5. Indexing – Vectors + semantic metadata live in local SQLite

6. Codemap – A lightweight `pampa.codemap.json` commits to git so context follows the repo

7. Serving – An MCP server exposes intelligent search and retrieval tools

Any MCP-compatible agent (Cursor, Claude, etc.) can now search with natural language, get instant responses for learned patterns, and stay synchronized – without scanning the entire tree.

πŸ€– For AI Agents & Humans

> πŸ€– If you're an AI agent: Read the complete setup guide for agents β†’

> or

> πŸ‘€ If you're human: Share the agent setup guide with your AI assistant to automatically configure PAMPA!

πŸ“š Table of Contents

🧠 Semantic Features

🏷️ Automatic Semantic Tagging

PAMPA automatically extracts semantic tags from your code without any special comments:

javascript
// File: app/Services/Payment/StripeService.php
function createCheckoutSession() { ... }

Automatic tags: `["stripe", "service", "payment", "checkout", "session", "create"]`

The system learns from successful searches and provides instant responses:

bash
# First search (vector search)
"stripe payment session" β†’ 0.9148 similarity

# System automatically learns and caches this pattern
# Next similar searches are instant:
"create stripe session" β†’ instant response (cached)
"stripe checkout session" β†’ instant response (cached)

πŸ“ˆ Adaptive Learning System

  • Automatic Learning: Saves successful searches (>80% similarity) as intentions
  • Query Normalization: Understands variations: `"create"` = `"crear"`, `"session"` = `"sesion"`
  • Pattern Recognition: Groups similar queries: `"[PROVIDER] payment session"`

🏷️ Optional @pampa-comments (Complementary)

Enhance search precision with optional JSDoc-style comments:

javascript
/**
 * @pampa-tags: stripe-checkout, payment-processing, e-commerce-integration
 * @pampa-intent: create secure stripe checkout session for payments
 * @pampa-description: Main function for handling checkout sessions with validation
 */
async function createStripeCheckoutSession(sessionData) {
	// Your code here...
}

Benefits:

  • +21% better precision when present
  • Perfect scores (1.0) when query matches intent exactly
  • Fully optional: Code without comments works automatically
  • Retrocompatible: Existing codebases work without changes

πŸ“Š Search Performance Results

Search TypeWithout @pampaWith @pampaImprovement
Domain-specific0.73310.8874+21%
Intent matching~0.61.0000+67%
General search0.6-0.80.8-1.0+32-85%

πŸ“ Supported Languages

PAMPA can index and search code in several languages out of the box:

  • JavaScript / TypeScript (`.js`, `.ts`, `.tsx`, `.jsx`)
  • PHP (`.php`)
  • Python (`.py`)
  • Go (`.go`)
  • Java (`.java`)

1. Configure your MCP client

Claude Desktop

Add to your Claude Desktop config (`~/Library/Application Support/Claude/claude_desktop_config.json` on macOS):

json
{
	"mcpServers": {
		"pampa": {
			"command": "npx",
			"args": ["-y", "pampa", "mcp"]
		}
	}
}

Optional: Add `"--debug"` to args for detailed logging: `["-y", "pampa", "mcp", "--debug"]`

Cursor

Configure Cursor by creating or editing the `mcp.json` file in your configuration directory:

json
{
	"mcpServers": {
		"pampa": {
			"command": "npx",
			"args": ["-y", "pampa", "mcp"]
		}
	}
}

2. Let your AI agent handle the indexing

Your AI agent should automatically:

  • Check if the project is indexed with `get_project_stats`
  • Index the project with `index_project` if needed
  • Keep it updated with `update_project` after changes

Need to index manually? See Direct CLI Usage section.

3. Install the usage rule for your agent

Additionally, install this rule in your application so it uses PAMPA effectively:

Copy the content from RULE_FOR_PAMPA_MCP.md into your agent or AI system instructions.

4. Ready! Your agent can now search code

Once configured, your AI agent can:

code
πŸ” Search: "authentication function"
πŸ“„ Get code: Use the SHA from search results
πŸ“Š Stats: Get project overview and statistics
πŸ”„ Update: Keep memory synchronized

πŸ’» Direct CLI Usage

For direct terminal usage or manual project indexing:

Install the CLI

bash
# Run without installing
npx pampa --help

# Or install globally (requires Node.js 20+)
npm install -g pampa

Index or update a project

bash
# Index current repository with the best available provider
npx pampa index

# Force the local CPU embedding model (no API keys required)
npx pampa index --provider transformers

# Re-embed after code changes
npx pampa update

# Inspect indexed stats at any time
npx pampa info

> Indexing writes `.pampa/` (SQLite database + chunk store) and `pampa.codemap.json`. Commit the codemap to git so teammates and CI re-use the same metadata.

CommandPurpose
`npx pampa index [path] [--provider X]`Create or refresh the full index at the provided path
`npx pampa update [path] [--provider X]`Force a full re-scan (helpful after large refactors)
`npx pampa watch [path] [--provider X]`Incrementally update the index as files change
`npx pampa search `Hybrid BM25 + vector search with optional scoped filters
`npx pampa context `Manage reusable context packs for search defaults
`npx pampa mcp`Start the MCP stdio server for editor/agent integrations

Search with scoped filters & ranking flags

`pampa search` supports the same filters used by MCP clients. Combine glob patterns, semantic tags, language filters, provider overrides, and ranking controls:

Flag / optionEffect
`--path_glob`Limit results to matching files (`"app/Services/**"`)
`--tags`Filter by codemap tags (`stripe`, `checkout`)
`--lang`Filter by language (`php`, `ts`, `py`)
`--provider`Override embedding provider for the query (`openai`, `transformers`)
`--reranker`Reorder top results with the Transformers cross-encoder (`off``transformers`)
`--hybrid` / `--bm25`Toggle reciprocal-rank fusion or the BM25 candidate stage (`on``off`)
`--symbol_boost`Toggle symbol-aware ranking boost that favors signature matches (`on``off`)
`-k, --limit`Cap returned results (defaults to 10)
bash
# Narrow to service files tagged stripe in PHP
npx pampa search "create checkout session" --path_glob "app/Services/**" --tags stripe --lang php

# Use OpenAI embeddings but keep hybrid fusion enabled
npx pampa search "payment intent status" --provider openai --hybrid on --bm25 on

# Reorder top candidates locally
npx pampa search "oauth middleware" --reranker transformers --limit 5

# Disable signature boosts for literal keyword hunts
npx pampa search "token validation" --symbol_boost off

> PAMPA extracts function signatures and lightweight call graphs with tree-sitter. When symbol boosts are enabled, queries that mention a specific method, class, or a directly connected helper will receive an extra scoring bump.

> When a context pack is active, the CLI prints the pack name before executing the search. Any explicit flag overrides the pack defaults.

Manage context packs

Store JSON packs in `.pampa/contextpacks/*.json` to capture reusable defaults:

jsonc
// .pampa/contextpacks/stripe-backend.json
{
	"name": "Stripe Backend",
	"description": "Scopes searches to the Stripe service layer",
	"path_glob": ["app/Services/**"],
	"tags": ["stripe"],
	"lang": ["php"],
	"reranker": "transformers",
	"hybrid": "off"
}
bash
# List packs and highlight the active one
npx pampa context list

# Inspect the full JSON definition
npx pampa context show stripe-backend

# Activate scoped defaults (flags still win if provided explicitly)
npx pampa context use stripe-backend

# Clear the active pack (use "none" or "clear")
npx pampa context use clear

MCP tip: The MCP tool `use_context_pack` mirrors the CLI. Agents can switch packs mid-session and every subsequent `search_code` call inherits those defaults until cleared.

Watch and incrementally re-index

bash
# Watch the repository with a 750β€―ms debounce and local embeddings
npx pampa watch --provider transformers --debounce 750

The watcher batches filesystem events, reuses the Merkle hash store in `.pampa/merkle.json`, and only re-embeds touched files. Press `Ctrl+C` to stop.

Run the synthetic benchmark harness

bash
npm run bench

The harness seeds a deterministic Laravel + TypeScript corpus and prints a summary table with Precision@1, MRR@5, and nDCG@10 for Base, Hybrid, and Hybrid+Cross-Encoder modes. Customise scenarios via flags or environment variables:

  • `npm run bench -- --hybrid=off` – run vector-only evaluation
  • `npm run bench -- --reranker=transformers` – force the cross-encoder
  • `PAMPA_BENCH_MODES=base,hybrid npm run bench` – limit to specific modes
  • `PAMPA_BENCH_BM25=off npm run bench` – disable BM25 candidate generation

Benchmark runs never download external models when `PAMPA_MOCK_RERANKER_TESTS=1` (enabled by default inside the harness).

An end-to-end context pack example lives in `examples/contextpacks/stripe-backend.json`.

🧠 Embedding Providers

PAMPA supports multiple providers for generating code embeddings:

ProviderCostPrivacyInstallation
Transformers.js🟒 Free🟒 Total`npm install @xenova/transformers`
Ollama🟒 Free🟒 TotalInstall Ollama + `npm install ollama`
OpenAIπŸ”΄ ~$0.10/1000 functionsπŸ”΄ NoneSet `OPENAI_API_KEY`
Cohere🟑 ~$0.05/1000 functionsπŸ”΄ NoneSet `COHERE_API_KEY` + `npm install cohere-ai`

Recommendation: Use Transformers.js for personal development (free and private) or OpenAI for maximum quality.

πŸ† Performance Analysis

PAMPA v1.12 uses a specialized architecture for semantic code search with measurable results.

πŸ“Š Performance Metrics

Synthetic Benchmark Results:

code
| Setting    | P@1   | MRR@5 | nDCG@10 |
| ---------- | ----- | ----- | ------- |
| Base       | 0.750 | 0.833 | 0.863   |
| Hybrid     | 0.875 | 0.917 | 0.934   |
| Hybrid+CE  | 1.000 | 0.958 | 0.967   |

🎯 Search Examples

bash
# Search for authentication functions
pampa search "user authentication"
β†’ AuthController::login, UserService::authenticate, etc.

# Search for payment processing
pampa search "payment processing"
β†’ PaymentService::process, CheckoutController::create, etc.

# Search with specific filters
pampa search "database operations" --lang php --path_glob "app/Models/**"
β†’ UserModel::save, OrderModel::find, etc.

**πŸ“ˆ Read Full Analysis β†’**

πŸš€ Architectural Advantages

1. Specialized Indexing - Persistent index with function-level granularity

2. Hybrid Search - BM25 + Vector + Cross-encoder reranking combination

3. Code Awareness - Symbol boosting, AST analysis, function signatures

4. Multi-Project - Native support for context across different codebases

Result: Optimized architecture for semantic code search with verifiable metrics.

πŸ—οΈ Architecture

code
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€ Repo (git) ─────────-──┐
β”‚ app/… src/… package.json etc.      β”‚
β”‚ pampa.codemap.json                 β”‚
β”‚ .pampa/chunks/*.gz(.enc)          β”‚
β”‚ .pampa/pampa.db (SQLite)           β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
          β–²       β–²
          β”‚ write β”‚ read
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚ indexer.js        β”‚   β”‚
β”‚ (pampa index)     β”‚   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β–²β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
          β”‚ store       β”‚ vector query
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚ gz fetch
β”‚ SQLite (local)     β”‚  β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β–²β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
          β”‚ read        β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
β”‚ mcp-server.js      β”‚β—„β”€β”˜
β”‚ (pampa mcp)        β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Key Components

LayerRoleTechnology
IndexerCuts code into semantic chunks, embeds, writes codemap and SQLitetree-sitter, openai@v4, sqlite3
CodemapGit-friendly JSON with {file, symbol, sha, lang} per chunkPlain JSON
Chunks dir.gz code bodies (or .gz.enc when encrypted) (lazy loading)gzip β†’ AES-256-GCM when enabled
SQLiteStores vectors and metadatasqlite3
MCP ServerExposes tools and resources over standard MCP protocol@modelcontextprotocol/sdk
LoggingDebug and error logging in project directoryFile-based logs

πŸ”§ Available MCP Tools

The MCP server exposes these tools that agents can use:

`search_code`

Search code semantically in the indexed project.

  • Parameters:
    • Database Location: `{path}/.pampa/pampa.db`
    • Returns: List of matching code chunks with similarity scores and SHAs

    `get_code_chunk`

    Get complete code of a specific chunk.

    • Parameters:
      • Chunk Location: `{path}/.pampa/chunks/{sha}.gz` or `{sha}.gz.enc`
      • Returns: Complete source code

      `index_project`

      Index a project from the agent.

      • Parameters:
        • Creates:
          • Effect: Updates database and codemap

          `update_project`

          πŸ”„ CRITICAL: Use this tool frequently to keep your AI memory current!

          Update project index after code changes (recommended workflow tool).

          • Parameters:
            • Updates:
              • When to use:
                • Effect: Keeps your AI agent's code memory synchronized with current state

                `get_project_stats`

                Get indexed project statistics.

                • Parameters:
                  • Database Location: `{path}/.pampa/pampa.db`
                  • Returns: Statistics by language and file

                  πŸ“Š Available MCP Resources

                  `pampa://codemap`

                  Access to the complete project code map.

                  `pampa://overview`

                  Summary of the project's main functions.

                  🎯 Available MCP Prompts

                  `analyze_code`

                  Template for analyzing found code with specific focus.

                  `find_similar_functions`

                  Template for finding existing similar functions.

                  πŸ” How Retrieval Works

                  • Vector search – Cosine similarity with advanced high-dimensional embeddings
                  • Summary fallback – If an agent sends an empty query, PAMPA returns top-level summaries so the agent understands the territory
                  • Chunk granularity – Default = function/method/class. Adjustable per language

                  πŸ“ Design Decisions

                  • Node only β†’ Devs run everything via `npx`, no Python, no Docker
                  • SQLite over HelixDB β†’ One local database for vectors and relations, no external dependencies
                  • Committed codemap β†’ Context travels with repo β†’ cloning works offline
                  • Chunk granularity β†’ Default = function/method/class. Adjustable per language
                  • Read-only by default β†’ Server only exposes read methods. Writing is done via CLI

                  🧩 Extending PAMPA

                  IdeaHint
                  More languagesInstall tree-sitter grammar and add it to `LANG_RULES`
                  Custom embeddingsExport `OPENAI_API_KEY` or switch OpenAI for any provider that returns `vector: number[]`
                  SecurityRun behind a reverse proxy with authentication
                  VS Code PluginPoint an MCP WebView client to your local server

                  πŸ” Encrypting the Chunk Store

                  PAMPA can encrypt chunk bodies at rest using AES-256-GCM. Configure it like this:

                  1. Export a 32-byte key in base64 or hex form:

                  bash
                  export PAMPA_ENCRYPTION_KEY="$(openssl rand -base64 32)"

                  2. Index with encryption enabled (skips plaintext writes even if stale files exist):

                  bash
                  npx pampa index --encrypt on

                  Without `--encrypt`, PAMPA auto-encrypts when the environment key is present. Use `--encrypt off` to force plaintext (e.g., for debugging).

                  3. All new chunks are stored as `.gz.enc` and require the same key for CLI or MCP chunk retrieval. Missing or corrupt keys surface clear errors instead of leaking data.

                  Existing plaintext archives remain readable, so you can enable encryption incrementally or rotate keys by re-indexing.

                  🀝 Contributing

                  1. Fork β†’ create feature branch (`feat/...`)

                  2. Run `npm test` (coming soon) & `npx pampa index` before PR

                  3. Open PR with context: why + screenshots/logs

                  All discussions on GitHub Issues.

                  πŸ“œ License

                  MIT – do whatever you want, just keep the copyright.

                  Happy hacking! πŸ’™


                  πŸ‡¦πŸ‡· Made with ❀️ in Argentina | πŸ‡¦πŸ‡· Hecho con ❀️ en Argentina

                  Frequently asked questions

                  What is pampa?

                  pampa is Protocol for Augmented Memory of Project Artifacts (MCP compatible)

                  How do I install pampa?

                  Open the GitHub repository and follow its README. Most MCP servers are added to your client's MCP config, then called by your agent.

                  Is pampa open source?

                  Yes β€” it is hosted on GitHub at https://github.com/tecnomanu/pampa and has 30 stars.

                  Related MCP tools

                  Run your own MCP server? See who uses it and what to fix.

                  Measure it with TrackMCP