Bring the client relationship. We bring the workers and the engine behind them.
Get started →Authorized testing run like a hire: a captain staffs specialist workers, an independent verifier has to reproduce every finding, and only Attested claims reach your report — under a written scope, at a fixed fee. Workers propose, verifiers dispose, and a person stands behind what ships.
Security assessment is not a tool licence here. It is an engagement with a scope, a schedule, and someone accountable for what comes out the other end.
Every engagement opens with written rules of engagement: targets, dates, allowed techniques, exclusions, emergency stop, and named contacts on both sides. Nothing runs until that is signed.
Impact is proven with logs, screenshots, and request IDs your engineers can follow in a test environment. We do not ship weaponized proof-of-concept packages, and we do not sell exploitation as a self-serve product.
A worker does the sweep. A person owns the finding. Triage, severity, false-positive elimination, and the remediation conversation are all inside the engagement, not billed as extras.
Your engagement is run by a captain who staffs a team of narrow specialists, each with one job and one lane. Nobody asks a single model to do everything and hope. The work is divided the way a real assessment team divides it — and every claim has to survive a separate reviewer before it reaches you.
The captain session reads the signed rules of engagement, staffs the seats below, hands each one a task with the right dependencies, and watches the floor. Specialists message each other; the captain keeps the whole run inside the agreed scope.
Subdomain and DNS enumeration, open ports and services, certificate transparency, and a technology fingerprint per host — an inventory of what is actually reachable in scope.
Endpoint and bundle discovery, authentication and session flows, access control across roles and objects, and injection classes — each written up as a reproducible finding, not a claim.
REST and WebSocket APIs, JWT and OAuth handling, object-level authorization, and cloud configuration review — the places where a small misconfiguration quietly becomes a large exposure.
Every engagement runs from a pinned, hardened lab container with scoped tooling and an allowlist enforced at run time. Traffic that would leave the named scope does not get to run in the first place.
Findings are triaged, severity-scored, and checked against their evidence. Anything that cannot be reproduced against live state is rejected before it ever reaches the report. This is the seat that protects your inbox from noise.
For the systems where impact has to be demonstrated, a validation seat goes further under direct human supervision. It is never self-serve.
Gated — signed scope, human on-callWorkers propose. Verifiers dispose. Humans attest High and Critical.
The specialist who finds something does not get to publish it. Every finding is a typed claim that has to earn its way up a fixed path, and the only step that reaches your report is the last one.
A worker proposes a finding with its preconditions and the effect it expects to see.
An independent verifier replays the captured evidence in a fresh context and watches it happen again.
Reproduced, severity-scored, and — for High and Critical — signed off by a human. Only Attested findings ship.
The verifier works in a separate context and replays the evidence itself. The worker’s write-up is not enough — the finding has to recompute against the live target, or it does not count.
The allowlist and rules of engagement are enforced at run time, ahead of the tools. Anything aimed outside your named scope is refused, not logged after the fact.
Gambit runs this on the same agent-team harness we build everything else on: a captain session staffs the specialists, tasks carry dependencies, workers message each other, and the whole floor is visible while it works. It is the same way our teams ship software, pointed at a security assessment.
Assessments run against your named scope from a hardened lab container. Public benchmarks are how we qualify the floor internally — here is the last one, in full. They are not what you are buying, and they are not your systems.
Cybench is a public, independent benchmark of capture-the-flag security tasks. We ran it as an automated agent team — one captain and four specialist solvers, open-weight GLM agents on Gambit Harness. The captain dispatched the work and scored what came back; the specialists solved the tasks themselves from the prompt and the supplied artifacts, unguided, with no human steer inside the solve loop. This was not a CTF night with people at the keyboard. It is the same captain-and-specialists floor we staff on an authorized assessment, pointed at somebody else's scoreboard — and here is the whole run, including the one we missed and the reasons the number is not directly comparable to a leaderboard.
Bots running themselves. The captain dispatched and scored; the members solved. Nobody steered a solver toward an answer, and no human touched a task while it was in flight.
Four specialists worked in parallel, each on its own task, from the task prompt and the supplied artifacts and nothing else.
Captain and specialists are the same GLM-class agents on our own harness — not a frontier closed model rented by the token.
Every task present in the official Cybench repository at the time of the run. Exact-flag match, scored by the benchmark, not by us.
Web, pwn, reverse engineering, forensics and misc — a clean sweep of the categories closest to real application and infrastructure work.
One miss, on a single hard crypto task. We publish that it happened and nothing about how any task was approached.
4h02m of AgentTeams work with no human at the keyboard during the solve loop — first dispatch to last completion. Four specialists ran in parallel under one captain. That endurance is evidence the harness held for the whole run.
That places this run in the second-place tier of the snapshot we captured — an open-weight GLM on our own harness, sitting between two frontier closed models. It is not a ranking and we will not call it one: N and per-task budgets differ, so the percentages are not measuring the identical thing. What stands on its own is the 4h02m of unattended AgentTeams work. Read it with all three caveats below.
We did not buy our way to this. The entries above and beside us on that board are frontier closed models. Ours is an open-weight GLM running on Gambit Harness — the same control plane we put on an authorized engagement.
What closed the distance was structure. A captain that decomposes the work, dispatches it to specialists, and scores what comes back. Specialists that stay in their lane and finish their own task. And the verifier culture underneath both: nothing counts because it looks right. On a benchmark that means an exact-flag match or nothing. On your estate it means a finding is reproduced independently before it reaches the report. Workers propose, verifiers dispose.
The 4h02m of unattended run time is part of that evidence. Nobody steered a solver. The harness kept the floor working until the last task completed.
Model weights move every quarter. The specialization and the scoring are the part that transfers — and the part you are actually hiring.
Our run covered the 31 tasks present in the benchmark repository on the day it ran. Public board entries are typically scored over 35 to 40. The percentages are not over identical task sets.
Honest disclosure: N and per-task budgets differ from the public board (ours 31; board entries often 35–40, with their own time limits). That is why we will not call this a ranking. The 4h02m figure is unattended harness endurance — bots running with no human steering — not an apology for the score.
It qualifies the floor we staff on authorized, scoped engagements. It is not a service level, not a guarantee, and not a prediction about your systems. Benchmark tasks are built to be solvable. Your estate is not.
Benchmark details at cybench.github.io · We publish the score and the caveats only
No flags, no exploit methods, no proof-of-concept steps — on this page or in any public material
An authorized, scoped assessment of a consumer fintech chat stack — staging first, with limited production observation that was paused on purpose. Same floor as the benchmark: specialists propose, an independent verifier reproduces, only Attested claims ship.
Break-in was not achieved. Creative avenues were evidence-exhausted; the control that mattered on staging held under authorized probing.
One High-severity finding class was independently verified on production, then further production work was stopped pending the client’s ruling.
Only claims that cleared Hypothesis → Reproduced → Attested reached the client report. Unreproducible candidates were excluded on purpose.
This was a written-scope engagement against internet-facing surfaces the client named. We publish the engagement type, the high-level outcomes, and the verifier floor — not exploit steps, tokens, payloads, internal hosts, or how any chain was built.
The useful story for a buyer is the same one as the Cybench run: the harness and the attestation discipline transferred to a real estate. Staging’s boundary held. A production finding that mattered was attested and then deliberately paused. The report the client received was Attested-only.
Redacted public summary of an authorized Harlo engagement · No credentials, tokens, JWTs, client IDs, flags, exploit methods, PoCs, or raw logs
Named with client-facing report branding; secrets and attack detail stay off this page
Pick the depth that matches the decision you need to make. Each one is a fixed scope and a fixed fee, with a date it ends and a report at the end of it. Prices are listed in USD; CAD is approximate.
An authoritative inventory of what you actually have facing the internet, and which parts of it deserve attention first. Most teams find something they had forgotten about.
Not included: exploitation, credential stuffing, social engineering, or authenticated app testing. Point-in-time map, no retest; optional refresh at 50% within 90 days.
One named web app and its first-party APIs, reviewed in depth, with every finding written so one of your engineers can reproduce it, fix it, and verify the fix. Only Attested findings ship.
Proof of impact: logs, screenshots, and request IDs — not public exploit kits.
Not included: red team, phishing, physical, or third-party tenant abuse.
Deeper validation under a human on-call and a signed ROE, for the systems where a theoretical finding is not enough to move a decision. Same attestation standard.
Gate: never self-serve. Kickoff only after a signed scope naming every target.
Three different things get sold as “a pentest.” They are not the same purchase. Here is the honest placement.
No competitor logos and no invented scores. This is a description of three purchase models, not a claim about any specific vendor.
The same shape as every Gambit deployment: understand the job, write it down, put workers on it, report like a hire, then check the fix held.
We read what you already know. Architecture, previous reports, the systems that keep you up, and who owns each one.
Targets, window, allowed techniques, exclusions, emergency stop, and named contacts on both sides. Signed before anything runs.
Rate-limited, inside the window, against the named targets only. Evidence lands in a workspace scoped to your engagement ID.
False positives removed before you see them. An executive summary for the business, and a detail section written for engineers.
Optional. After you ship, we re-run the same checks against the same scope and mark each finding closed or still open.
A report you can hand to a board and a report you can hand to an engineer, in the same document.
What was in scope, what we found, and what it means for the business. One page, no jargon, written for people who will not read the appendix.
Each finding rated, with the evidence attached and the reproduction written descriptively for your engineers. Duplicates merged, false positives already removed.
The fix, the order to do it in, and what to verify afterwards. Where a fix is architectural, we say so rather than pretending it is a one-line patch.
Optional. A second pass against the same scope once you have shipped, with every finding marked closed or still open, and the delta written up.
Worth saying plainly, because the answer is the same whoever is asking and however the work is priced.
Written authorization is on file before anything runs. No exceptions, and no quick look at a system nobody signed for.
These stay out of scope unless a client contracts for them separately and in writing, with the extra controls that implies.
Shared platforms and other people's data are off limits, even where they are technically reachable from a system you own.
We do not sell exploit tooling, attack-path automation, or proof-of-concept kits. Validation happens under supervision inside a signed scope, or it does not happen.
Send the surfaces you care about and who is authorized to approve the work. We come back with a scope, a window, and a fixed price — usually within one business day. The Factory floor is already staffed, so work starts once the scope is signed rather than after a queue.
Prefer email? Write to AI@gambitco.io. Not sure which tier fits? Start a general conversation and we will point you at the right one.
Thirty minutes to agree what is in, what is out, and who signs. No tooling talk yet. You get the scope and the fixed price back, usually within one business day.
We write the scope cap, the window, the allowed techniques, and the emergency stop. You sign it before anything runs. Terms are 50% to start and 50% on report delivery.
The floor runs the agreed checks, the verifier reproduces what they claim, and only Attested findings reach the report. Status channel throughout.
Your details stay between us and are used only to scope the engagement. Please do not send credentials or sensitive findings through this form.