Honest assessment of what these skills cover vs. what real practitioners need.
~85–90% coverage of what an experienced practitioner would reach for during the OSINT/recon phase of an authorized external red-team engagement.
~35–45% coverage of what a full external red-team operator does in their job (because most red-team work is exploitation + post-exploit, intentionally out of scope).
The six depth skills extend the core recon pair past single-target discovery into enterprise attack-surface reasoning — the areas earlier versions were thinnest on:
org-attack-surface— legal-entity → owned-footprint attribution (subsidiaries, dark netblocks, ASN scope). Closes the "whose assets are these, really?" gap the core pair left to the operator.email-domain-security— a defensible spoofability verdict (not just a record dump) + SPF supply-chain analysis.exposure-risk-quantification— turns findings into a 0–100 + A–F score and a $-denominated board deliverable. Lifts the Reporting phase.continuous-exposure-monitoring— re-scan/diff + CTI chatter + FP discipline. Lifts the previously-thin Re-test / continuous-monitoring phase.cloud-saas-exposure— ownership-gated bucket triage, offline AWS account-ID decode, dependency-confusion confirmation.identity-provider-recon— tenant/federation mapping + a disciplined pre-auth user-enumeration oracle with a hard boundary.
All six stay strictly passive/read-only; each carries an explicit boundary that stops it short of active exploitation.
| Archetype | Coverage of their needs | Why |
|---|---|---|
| Pure OSINT analyst | ~90% | Skills are built for this. |
| External attack-surface analyst (CyCognito-style) | ~90% | Direct overlap with the methodology; the org-grade skills add entity→footprint attribution, $-denominated risk scoring, and continuous re-scan/diff — the CyCognito-class core. |
| Bug bounty hunter | ~75–80% | Strong on recon; thin on exploit techniques. |
| Threat intel investigator | ~75% | RU/CN pivots, attribution discipline, malware basics, plus infrastructure-tracking-over-time and adversary-chatter/ransomware-leak monitoring (continuous-exposure-monitoring). |
| External red teamer (recon phase) | ~85–90% | The OSINT phase is well-covered. |
| External red teamer (full engagement) | ~35–45% | Recon is ~30–40% of a full engagement; rest (exploitation, post-exploit, lateral, reporting) is mostly out of scope. |
| Internal red teamer (assumed-breach) | ~10% | Almost entirely out of scope. |
| Adversary emulation / TTP-driven | ~25% | Threat-actor section exists; specific TTP playbooks per APT don't. |
| Physical pentester | ~25% | Sat imagery + LinkedIn intel cover scouting; physical execution doesn't. |
| Social engineer | ~50% | Pretext development covered; payload crafting + voice tradecraft not. |
| Purple teamer | ~30% | No SOC-coordination guidance. |
| Phase | Coverage |
|---|---|
| Pre-engagement (RoE, scoping, NDAs, SOW, pricing) | ~10% |
| External OSINT / passive recon | ~85–90% |
| External active recon (light probing) | ~75–85% |
| Phishing payload crafting + delivery | 0% (out of scope) |
| Initial access (exploit execution) | ~5% (we identify, don't exploit) |
| Foothold / persistence | 0% (out of scope) |
| Privilege escalation (local + AD) | 0% (out of scope) |
| Lateral movement | 0% (out of scope) |
| C2 infrastructure | 0% (out of scope) |
| AV/EDR evasion | 0% (out of scope) |
| Domain dominance | 0% (out of scope) |
| Data exfiltration tradecraft | 0% (out of scope) |
| Cleanup / artifact removal | 0% (out of scope) |
| Reporting (technical + exec + board $) | ~85% |
| Disclosure / vendor coordination | ~60% |
| Re-test / continuous monitoring | ~75% (continuous-exposure-monitoring: re-scan/diff loop + CTI chatter + FP discipline + alert outbox) |
| Purple-team / SOC-coordination | 0% |
| Lessons-learned / engagement retrospective | ~20% |
- Active exploitation, post-exploitation, malware — operational tradecraft, different domain, safety posture concerns.
- C2 frameworks, AV/EDR evasion — operational tradecraft, large body of separate knowledge.
- AD attacks, BloodHound, Kerberos — internal recon, not external.
- Specific client-portal report formats — too company-specific to template usefully.
- Pricing, NDA, SOW templates — business operations, not technical.
- Real PII / breach corpus content — privacy + opsec.
The repo ships 56 self-test prompts (tests/smoke-test-prompts.md) covering the major capability areas — including Tier 4 prompts (34–48) for the six org-grade depth skills and hard-boundary refusal tests (B3–B8).
| Run | PASS | PARTIAL | FAIL | Grade |
|---|---|---|---|---|
| v2.0 (initial) | 1 | 9 | 22 | C |
| v2.1 (32 prompts) | 31 | 1 | 0 | A |
| current (56 prompts, 8 skills) | 56 | 0 | 0 | A+ |
The current run was executed against fresh sessions (each blind to the expected-behavior column, auto-routing from skill frontmatter): zero fabricated endpoints/regexes/sections, and every hard-boundary prompt correctly refused with a real section cite.
The smoke-test number (100% PASS on 56 prompts) is Claude grading itself on tests Claude designed. It's a useful signal for tracking gaps but not an objective measure of real-world coverage. A real practitioner would find more gaps. Treat it as "the skills now answer the obvious questions"; non-obvious questions may need a follow-on iteration.
If a senior offensive consultant reviewed v2.1 and stayed within OSINT scope, here's what they'd flag as still missing:
- Specific tool-chaining recipes — "use spiderfoot → export CSV → maltego transforms → asset graph" workflows. We name tools; we don't compose them step-by-step.
- Recon-ng / SpiderFoot / Maltego module-by-module configuration — these are full ecosystems; we treat them as pointers.
- Custom Burp Suite / OWASP ZAP setup for engagements — the "configure your active proxy for an engagement" guide.
- OPSEC infrastructure as code — Terraform/Ansible to spin up clean engagement infrastructure (proxy stacks, redirectors).
- Sector-specific deep dives — §47 is a starting point, not a deep dive (real healthcare RT specialists know HL7 trafficking like a second language).
- Adversary-emulation playbooks per APT — "to simulate APT29's external recon, use these specific tools/techniques."
Continuous-monitoring orchestration — daily diff scripts, alert pipelines, false-positive tuning.Shipped in thecontinuous-exposure-monitoringskill (baseline→re-scan→diff→threshold-alert loop, CTI chatter, finding-lifecycle FP discipline, durable alert outbox).- Multi-tenant engagement workflow — how an MSSP runs 30 concurrent ASM engagements without crossing wires.
- Client-specific report styling — every Big-4 consultancy has their own template.
- Tool failure recovery — when Shodan rate-limits during a critical phase, what's plan B/C/D?
These would push coverage to ~95% of OSINT-phase work. Each would add 200–500 lines and approach the limits of what a single skill can usefully encode.
| Phase | Status | Description |
|---|---|---|
| v1.0 | ✅ Done | Original framework |
| v2.0 | ✅ Done | External-red-team posture rewrite |
| v2.1 | ✅ Done | Comprehensive expansion of the core pair |
| Org-grade depth | ✅ Current | Six new skills — org-attack-surface, email-domain-security, exposure-risk-quantification, continuous-exposure-monitoring, cloud-saas-exposure, identity-provider-recon; offensive-osint v2.2 (80-pattern catalog + APK); osint-methodology v2.3 (6-stage pipeline) |
| Next | 🔜 | Multi-tenant / MSSP engagement workflow · Burp extension recipes · adversary-emulation playbooks per APT |
| v3.0 | 🔜 | Plugin manifest for one-click Claude Code install + optional MCP server companion |
For "external OSINT for authorized red-team operations": ~85–90% coverage of what an experienced practitioner reaches for. For "everything a full red-team operator does in their job": ~35–45% — the gap is mostly intentional (out of scope).
The skills are production-ready for OSINT-phase work. They are not a replacement for a senior red teamer on a full engagement.