Get Started Ask Gambit
Security assessments

Authorized testing, run like a hire.

A named owner. An independent verifier. A fixed fee.

Authorized testing run like a hire: a captain staffs specialist workers, an independent verifier has to reproduce every finding, and only Attested claims reach your report — under a written scope, at a fixed fee. Workers propose, verifiers dispose, and a person stands behind what ships.

Rules of engagement Signed before any traffic
In scopeNamed hosts, IP ranges, and applications. If it is not on the list, it is not touched.
WindowAgreed dates and hours. Active checks are rate-limited so your monitoring stays readable.
TechniquesPassive, active non-destructive, authenticated. Each one listed and agreed, never assumed.
Out of scopeDenial of service. Social engineering. Physical entry. Third-party tenants and shared platforms.
Emergency stopA named on-call on both sides and one channel that reaches a human at any hour.
How we work

Authorized work, or no work.

Security assessment is not a tool licence here. It is an engagement with a scope, a schedule, and someone accountable for what comes out the other end.

Scope before tooling

Every engagement opens with written rules of engagement: targets, dates, allowed techniques, exclusions, emergency stop, and named contacts on both sides. Nothing runs until that is signed.

Evidence, not exploit kits

Impact is proven with logs, screenshots, and request IDs your engineers can follow in a test environment. We do not ship weaponized proof-of-concept packages, and we do not sell exploitation as a self-serve product.

Owned end to end

A worker does the sweep. A person owns the finding. Triage, severity, false-positive elimination, and the remediation conversation are all inside the engagement, not billed as extras.

How the assessment floor is staffed

A floor of specialists, not one chatty bot.

Your engagement is run by a captain who staffs a team of narrow specialists, each with one job and one lane. Nobody asks a single model to do everything and hope. The work is divided the way a real assessment team divides it — and every claim has to survive a separate reviewer before it reaches you.

Captain

Plans the engagement and holds the scope.

The captain session reads the signed rules of engagement, staffs the seats below, hands each one a task with the right dependencies, and watches the floor. Specialists message each other; the captain keeps the whole run inside the agreed scope.

Recon

Maps the surface.

Subdomain and DNS enumeration, open ports and services, certificate transparency, and a technology fingerprint per host — an inventory of what is actually reachable in scope.

Web app

Tests the application.

Endpoint and bundle discovery, authentication and session flows, access control across roles and objects, and injection classes — each written up as a reproducible finding, not a claim.

API & cloud config

Checks the interfaces and settings.

REST and WebSocket APIs, JWT and OAuth handling, object-level authorization, and cloud configuration review — the places where a small misconfiguration quietly becomes a large exposure.

Infra & lab

Runs the hardened lab.

Every engagement runs from a pinned, hardened lab container with scoped tooling and an allowlist enforced at run time. Traffic that would leave the named scope does not get to run in the first place.

Report

Triages, scores, and kills false positives.

Findings are triaged, severity-scored, and checked against their evidence. Anything that cannot be reproduced against live state is rejected before it ever reaches the report. This is the seat that protects your inbox from noise.

Deeper validation

Only when a theoretical finding is not enough.

For the systems where impact has to be demonstrated, a validation seat goes further under direct human supervision. It is never self-serve.

Gated — signed scope, human on-call
The control plane

Workers propose. Verifiers dispose. Humans attest High and Critical.

The specialist who finds something does not get to publish it. Every finding is a typed claim that has to earn its way up a fixed path, and the only step that reaches your report is the last one.

Stage 1

Hypothesis

A worker proposes a finding with its preconditions and the effect it expects to see.

Stage 2

Reproduced

An independent verifier replays the captured evidence in a fresh context and watches it happen again.

Stage 3 · delivered

Attested

Reproduced, severity-scored, and — for High and Critical — signed off by a human. Only Attested findings ship.

Independent verification

The verifier works in a separate context and replays the evidence itself. The worker’s write-up is not enough — the finding has to recompute against the live target, or it does not count.

Scope gate before any tooling

The allowlist and rules of engagement are enforced at run time, ahead of the tools. Anything aimed outside your named scope is refused, not logged after the fact.

On our own harness

Run on the Gambit Factory harness.

Gambit runs this on the same agent-team harness we build everything else on: a captain session staffs the specialists, tasks carry dependencies, workers message each other, and the whole floor is visible while it works. It is the same way our teams ship software, pointed at a security assessment.

Where it runs

Against your scope, from a lab.

Assessments run against your named scope from a hardened lab container. Public benchmarks are how we qualify the floor internally — here is the last one, in full. They are not what you are buying, and they are not your systems.

Capability evidence

How we qualify the floor.

Cybench is a public, independent benchmark of capture-the-flag security tasks. We ran it as an automated agent team — one captain and four specialist solvers, open-weight GLM agents on Gambit Harness. The captain dispatched the work and scored what came back; the specialists solved the tasks themselves from the prompt and the supplied artifacts, unguided, with no human steer inside the solve loop. This was not a CTF night with people at the keyboard. It is the same captain-and-specialists floor we staff on an authorized assessment, pointed at somebody else's scoreboard — and here is the whole run, including the one we missed and the reasons the number is not directly comparable to a leaderboard.

How it ran
Automated agent team

Bots running themselves. The captain dispatched and scored; the members solved. Nobody steered a solver toward an answer, and no human touched a task while it was in flight.

Who ran it
1 captain · 4 specialist solvers

Four specialists worked in parallel, each on its own task, from the task prompt and the supplied artifacts and nothing else.

What it ran on
Open-weight GLM on Gambit Harness

Captain and specialists are the same GLM-class agents on our own harness — not a frontier closed model rented by the token.

Unguided exact-flag solve rate
30 / 31
96.8%

Every task present in the official Cybench repository at the time of the run. Exact-flag match, scored by the benchmark, not by us.

Non-crypto20 / 20

Web, pwn, reverse engineering, forensics and misc — a clean sweep of the categories closest to real application and infrastructure work.

Cryptography10 / 11

One miss, on a single hard crypto task. We publish that it happened and nothing about how any task was approached.

Unattended wall-clock4h02m

4h02m of AgentTeams work with no human at the keyboard during the solve loop — first dispatch to last completion. Four specialists ran in parallel under one captain. That endurance is evidence the harness held for the whole run.

Public board context — read with the caveats below
  • Mythos Preview100%@ 35 tasksFrontier closed model
  • This run · Gambit Harness96.8%@ 31 tasksOpen-weight GLM · automated agent team
  • Opus 4.796%@ 35 tasksFrontier closed model

That places this run in the second-place tier of the snapshot we captured — an open-weight GLM on our own harness, sitting between two frontier closed models. It is not a ranking and we will not call it one: N and per-task budgets differ, so the percentages are not measuring the identical thing. What stands on its own is the 4h02m of unattended AgentTeams work. Read it with all three caveats below.

What the number is actually evidence of

The result is the harness, not the model.

We did not buy our way to this. The entries above and beside us on that board are frontier closed models. Ours is an open-weight GLM running on Gambit Harness — the same control plane we put on an authorized engagement.

What closed the distance was structure. A captain that decomposes the work, dispatches it to specialists, and scores what comes back. Specialists that stay in their lane and finish their own task. And the verifier culture underneath both: nothing counts because it looks right. On a benchmark that means an exact-flag match or nothing. On your estate it means a finding is reproduced independently before it reaches the report. Workers propose, verifiers dispose.

The 4h02m of unattended run time is part of that evidence. Nobody steered a solver. The harness kept the floor working until the last task completed.

Model weights move every quarter. The specialization and the scoring are the part that transfers — and the part you are actually hiring.

Different N

Our run covered the 31 tasks present in the benchmark repository on the day it ran. Public board entries are typically scored over 35 to 40. The percentages are not over identical task sets.

Board comparison note

Honest disclosure: N and per-task budgets differ from the public board (ours 31; board entries often 35–40, with their own time limits). That is why we will not call this a ranking. The 4h02m figure is unattended harness endurance — bots running with no human steering — not an apology for the score.

What it is evidence of

It qualifies the floor we staff on authorized, scoped engagements. It is not a service level, not a guarantee, and not a prediction about your systems. Benchmark tasks are built to be solvable. Your estate is not.

Benchmark details at cybench.github.io · We publish the score and the caveats only
No flags, no exploit methods, no proof-of-concept steps — on this page or in any public material

Delivery evidence

Harlo: staging held. Findings attested.

An authorized, scoped assessment of a consumer fintech chat stack — staging first, with limited production observation that was paused on purpose. Same floor as the benchmark: specialists propose, an independent verifier reproduces, only Attested claims ship.

Authorized & scoped App & API / staging-tier Staging break-in not achieved Production High paused
Staging outcome
Gate held

Break-in was not achieved. Creative avenues were evidence-exhausted; the control that mattered on staging held under authorized probing.

Production posture
High class — paused

One High-severity finding class was independently verified on production, then further production work was stopped pending the client’s ruling.

Attested findings
17 published

Only claims that cleared Hypothesis → Reproduced → Attested reached the client report. Unreproducible candidates were excluded on purpose.

What we are willing to say in public.

  • Hypothesis
  • Reproduced
  • Attested

This was a written-scope engagement against internet-facing surfaces the client named. We publish the engagement type, the high-level outcomes, and the verifier floor — not exploit steps, tokens, payloads, internal hosts, or how any chain was built.

The useful story for a buyer is the same one as the Cybench run: the harness and the attestation discipline transferred to a real estate. Staging’s boundary held. A production finding that mattered was attested and then deliberately paused. The report the client received was Attested-only.

Redacted public summary of an authorized Harlo engagement · No credentials, tokens, JWTs, client IDs, flags, exploit methods, PoCs, or raw logs
Named with client-facing report branding; secrets and attack detail stay off this page

The offerings

Three ways to start.

Pick the depth that matches the decision you need to make. Each one is a fixed scope and a fixed fee, with a date it ends and a report at the end of it. Prices are listed in USD; CAD is approximate.

Surface Map
About 2 weeks
Recon and inventory
USD 4,800≈ CAD 6,500

An authoritative inventory of what you actually have facing the internet, and which parts of it deserve attention first. Most teams find something they had forgotten about.

Scope capUp to 25 hostnames or IPs, or one primary domain plus its discovered subdomains. Overage +USD 150/host or a custom quote.
Includes
  • Subdomain and DNS map
  • Technology fingerprint per host
  • Open ports and services table
  • Prioritized watchlist, ranked by exposure
  • Short executive note

Not included: exploitation, credential stuffing, social engineering, or authenticated app testing. Point-in-time map, no retest; optional refresh at 50% within 90 days.

Supervised Full Assessment
Scoped
Human oversight, explicit ROE
From USD 28,000≈ CAD 38,000

Deeper validation under a human on-call and a signed ROE, for the systems where a theoretical finding is not enough to move a decision. Same attestation standard.

Scope capScoped per engagement. The “from” price is the floor for one mid-complexity product surface; fixed after the scoping call.
Includes
  • Everything in App & API for the agreed surfaces
  • Supervised deeper-validation seats as needed
  • Your on-call available throughout the window
  • Coverage statement: what was and was not tested

Gate: never self-serve. Kickoff only after a signed scope naming every target.

Compared to what

Where this sits.

Three different things get sold as “a pentest.” They are not the same purchase. Here is the honest placement.

Vulnerability scanners

Cheap, automated, unverified

  • Signature and version matching at speed
  • Output is a candidate list, not findings — triage lands on your team
  • No access-control or business-logic reasoning
  • Good for continuous hygiene, weak for a decision
Day-rate boutiques

Skilled humans, open-ended meter

  • Typically priced by the day, so the total moves with the calendar
  • Scope tends to widen mid-engagement
  • Real expertise, but usually a queue before anyone starts
  • Quality tracks whichever tester you happen to get
Gambit assessment floor

Verified findings at a fixed fee

  • A captain staffs named specialist seats, not one general-purpose bot
  • An independent verifier has to reproduce every finding before it can ship
  • Only Attested claims reach the report; High and Critical are human-attested
  • Fixed fee with a hard scope cap and an end date, written before any traffic

No competitor logos and no invented scores. This is a description of three purchase models, not a claim about any specific vendor.

How it runs

Five steps, and you know where you are in all of them.

The same shape as every Gambit deployment: understand the job, write it down, put workers on it, report like a hire, then check the fix held.

01 · Discover

Understand the estate

We read what you already know. Architecture, previous reports, the systems that keep you up, and who owns each one.

02 · Build

Scope and ROE in writing

Targets, window, allowed techniques, exclusions, emergency stop, and named contacts on both sides. Signed before anything runs.

03 · Deploy

Workers run the agreed checks

Rate-limited, inside the window, against the named targets only. Evidence lands in a workspace scoped to your engagement ID.

04 · Report

Triaged, rated, de-duplicated

False positives removed before you see them. An executive summary for the business, and a detail section written for engineers.

05 · Retest

Confirm the fix held

Optional. After you ship, we re-run the same checks against the same scope and mark each finding closed or still open.

Evidence redacted before sharing · Scope allowlist enforced at run time · Nothing outside the named scope
Deliverables

What actually lands in your inbox.

A report you can hand to a board and a report you can hand to an engineer, in the same document.

01

Executive summary

What was in scope, what we found, and what it means for the business. One page, no jargon, written for people who will not read the appendix.

02

Findings with severity and evidence

Each finding rated, with the evidence attached and the reproduction written descriptively for your engineers. Duplicates merged, false positives already removed.

03

Remediation guidance

The fix, the order to do it in, and what to verify afterwards. Where a fix is architectural, we say so rather than pretending it is a one-line patch.

04

Retest window

Optional. A second pass against the same scope once you have shipped, with every finding marked closed or still open, and the delta written up.

Gambit · Findings viewIllustrative
14
Findings reported
2
High severity
31
False positives removed
9
Closed at retest
FindingSurfaceSeverityEvidenceStatus
Object-level access control gapOrders APIHighRequest IDsClosed
Session not invalidated on password changeAuth serviceHighSession logReported
Verbose error exposes stack tracesWeb appMediumScreenshotsClosed
Missing transport and content headersEdge configLowHeader dumpReported
Illustrative view of a report in progress. Not a customer engagement.
Boundaries

What we will not do.

Worth saying plainly, because the answer is the same whoever is asking and however the work is priced.

  • No unauthorized testing

    Written authorization is on file before anything runs. No exceptions, and no quick look at a system nobody signed for.

  • No denial of service, social engineering, or physical entry

    These stay out of scope unless a client contracts for them separately and in writing, with the extra controls that implies.

  • No third-party tenants

    Shared platforms and other people's data are off limits, even where they are technically reachable from a system you own.

  • No self-serve exploitation

    We do not sell exploit tooling, attack-path automation, or proof-of-concept kits. Validation happens under supervision inside a signed scope, or it does not happen.

Book a scoping call

Tell us what you want assessed.

Send the surfaces you care about and who is authorized to approve the work. We come back with a scope, a window, and a fixed price — usually within one business day. The Factory floor is already staffed, so work starts once the scope is signed rather than after a queue.

Prefer email? Write to AI@gambitco.io. Not sure which tier fits? Start a general conversation and we will point you at the right one.

1

A scoping call

Thirty minutes to agree what is in, what is out, and who signs. No tooling talk yet. You get the scope and the fixed price back, usually within one business day.

2

Rules of engagement

We write the scope cap, the window, the allowed techniques, and the emergency stop. You sign it before anything runs. Terms are 50% to start and 50% on report delivery.

3

The engagement

The floor runs the agreed checks, the verifier reproduces what they claim, and only Attested findings reach the report. Status channel throughout.

Engagement enquiry — priced in one business day

Your details stay between us and are used only to scope the engagement. Please do not send credentials or sensitive findings through this form.