PTDby Gonka Labs

← All proposals

Proposal #82

Not yet claimed on PTD. Claim this proposal to add tasks and post your own updates.

Report from Gonka GitHub Discussions

Proposal #82: External Test Lab & Community DevNet — M2 Report

paranjkoparanjko · Updated 9/13/2026

View on GitHub ↗

This report is community-posted and maintained by the Core team on GitHub, not verified by PTD moderators. Independently confirm important claims.

Period: 2026-08-10 — 2026-09-09 · Proposal: #82 (discussion)

Executive summary

Month 2 moved the Lab from setup into active infrastructure, Host lifecycle testing, release qualification, and broker compatibility work.

The rented fleet reached nine GPU hosts, but the 9+ validated MLNodes online target remains Partial: some hosts were deliberately used for destructive JOIN / reset / restore tests, two new NVIDIA hosts were still being prepared at cutoff, and one AMD host remains experimental.

The second QA / infrastructure testing engineer started on September 7 after ~160 responses, 13 interviews, and 2 finalists.

Validation produced concrete results: independent JOIN / recovery testing, immutable DevShard candidate artifacts, the active DevShard v5 / Gonka v0.2.16 qualification campaign, broker compatibility tooling, and a Mainnet gateway v4 regression affecting Kimi-K2.6 non-stream requests (gonka-ai/gonka#1680). The bug was fixed upstream and successfully retested.

On September 9, repeated node5 JOIN testing caused a high-severity Community DevNet consensus halt. Recovery and investigation continued after cutoff in #143.

M2 targets

Target Status Result
9+ MLNodes across target regions Partial Nine GPU hosts rented; not all qualified / online simultaneously. Some capacity reserved for JOIN / recovery testing.
Regional layout Done devnet/regions.md
Smoke + regression catalogues Done smoke · regression
Second QA engineer Done ~160 responses → 13 interviews → 2 finalists → started Sep 7
Validation work Done / active Host lifecycle validation and DevShard v5 / Gonka v0.2.16 qualification active (#28, #49)
External Host JOIN In progress Public interface available; independent real-server lifecycle validation is not yet stable

Community DevNet

During M2 the Lab renewed the original fleet and added four GPU hosts, reaching nine GPU hosts plus one network-only host across the US, Finland, the Netherlands, the UK, and a staged host in Russia.

The scale KPI remains Partial because rented capacity is not the same as validated online MLNode capacity. Some hosts were repeatedly rebuilt for JOIN / recovery testing; two new NVIDIA hosts were still in preparation; the AMD RX 9070 XT host remains experimental. See regional layout and live state on gonka-dev.net.

The public JOIN path also advanced from coordinator-driven setup toward an independently usable Host lifecycle: network-derived bootstrap / Join Profile handling, ML qualification under the resolved release, retained diagnostics, backup / restore work, and safer failure reporting. Real-server validation remains active in #28 and #117.

External Testing Team

The second QA hiring round completed with ~160 responses, 13 interviews, 2 finalists, and one selected engineer starting September 7. The new engineer received no payments during M2.

The public M1 strategy is now complemented by smoke and regression catalogues. They expose stable test coverage without publishing campaign-specific credentials, private data, or security-sensitive procedures.

Validation and findings

The main active release-assurance campaign is #49 — DevShard v5 and Gonka v0.2.16 qualification. By cutoff, the Lab had prepared immutable candidate definitions, verifiable artifacts / attestations, separate core and DevShard identities, feature scenarios, and controlled rollout / comparison work. No completed readiness verdict is claimed yet.

Testing of the Coreteam scenarios published in feature has started. Qualification is in progress, no final release-readiness verdict has been issued.

Independent JOIN testing produced reproducible defects and regression inputs. Separately, gonka-ai/gonka#1680 identified a Mainnet gateway v4 regression where Kimi-K2.6 non-stream requests could fail with 502 nonce_finished=false while streaming and v3 remained healthy. The fix was merged into gateway v4 and successfully retested on mainnet-v0.2.15-v4.0.1.

M2 also published broker-side OpenAI compatibility artifacts: Chat Completions host-compat guidance / shim, a POST /v1/responses adapter, and installable skills packaging. See broker-compat/ and skills/.

Incident — GNK-LAB-2026-0001

On September 9, repeated independent JOIN testing for node5 halted block production on gonka-devnet-community; the last block was 306552 at 2026-09-09T18:59:22Z.

Preliminary evidence indicates a consensus-key mismatch during JOIN. With another validator offline, remaining voting power fell below the CometBFT quorum.

Impact: DevNet operations requiring new blocks and JOIN verification halted. No Mainnet impact or loss of funds was detected. Recovery / investigation: #143.

Spending

Invoice amounts are the source of truth; original invoice currency is preserved. Raw invoices remain private because they may contain personal, payment, account, or infrastructure data.

Budget line Monthly cap Spent (M2) Notes
DevNet machines $5,000 $1,964.67 + €299.00 Renewals + node5–node8; includes staged / experimental capacity
Burst GPU rental $6,500 $10.00 SubModel evaluation
Tooling and reporting $250 $0
External Testing Engineers $8,000 $3,000 QA work rewards; M1 already reported $2,000
Contingency $2,250 $0
Total $22,000 $4,974.67 + €299.00

Infrastructure breakdown

Date Provider Configuration Role Amount
08-17 DatabaseMart RTX Pro 2000 GPU VPS node4-ml renewal $119.00
08-17 DatabaseMart RTX A5000 dedicated node0 renewal $254.77
08-17 OneProvider Tesla T4 dedicated node1 renewal $331.63
08-17 OneProvider network-only dedicated node4 late-billed prorated service $34.33
08-26 DatabaseMart RTX A4000 dedicated node5, JOIN / recovery target $279.00
08-27 LeaderGPU RTX 4090 dedicated node2 renewal €299.00
08-27 GIGAGPU RTX 3090 dedicated node3 renewal $270.03
09-07 SubModel GPU credit burst evaluation $10.00
09-07 HOSTKEY RTX 2000 PRO dedicated node6, staged $231.02
09-07 GIGAGPU RX 9070 XT dedicated node8, experimental $229.32
09-07 OneProvider Tesla T4 dedicated node7, staged $215.57

Month 3 plan

  • Stabilize Host JOIN / backup / restore and publish the incident outcome.
  • Complete the 9+ validated MLNodes online target.
  • Continue DevShard and protocol updates qualification and publish scoped verdicts.
  • Add smoke automation and M3 incident / participant-onboarding artifacts.
  • Complete second-QA onboarding and split repeatable validation ownership.

Report from Gonka GitHub Discussions

Proposal #82: External Test Lab & Community DevNet — M1 Report

paranjkoparanjko · Updated 8/11/2026

View on GitHub ↗

This report is community-posted and maintained by the Core team on GitHub, not verified by PTD moderators. Independently confirm important claims.

Period: 2026-07-10 — 2026-08-09 · Proposal: #82 (discussion) · Next unlock: 2026-08-13

Workstream A — DevNet infrastructure (M1: Setup)

M1 requirement Status
DevNet design agreed Done — architecture note
≥5 nodes online Done — node0–node4 deployed, one provider each, chain gonka-devnet-community; live state on the status site
Infrastructure blockers documented Done — see below

Delivered beyond the M1 bar:

  • Deployment runbook 1.0.0-alpha.1 (net-deployment-runbook/) — reproducible baseline on testnet-0.2.14 (phased CLI: prepare → identities → qualify-ml → genesis → join → verify), governed 0.2.14 → 0.2.15 upgrade rehearsal, single-node stop/start/verify/reset, network reset contract. The M2 artifact "runbook sufficient to reproduce a node" (ROLE-JOIN.md) is already published.
  • Operator handoff flow — an external operator can qualify, deploy, and register a node without receiving coordinator secrets; the coordinator confirms on-chain registration and grants funding/ML permission.
  • Public observabilitystatus site (per-node cards from committed chain state, honest OFFLINE on reset) and Grafana dashboards: 24h network view, 7-day inference view, triage overview.
  • Authenticated accessapi.gonka-dev.net/v1: OpenAI-compatible, chain-accounted DevShard v4 gateway (v3/v4 approved via on-chain governance), automatic escrow rotation across epochs, API keys self-issued through a Telegram bot from a finite pool, no artificial per-key quotas.

All images and binaries are pinned by digest; DevNet keys are fully isolated from mainnet.

Blockers encountered (documented and resolved)

  • Blackwell GPU support. The RTX Pro 2000 ML host (compute capability 12.0) is not supported by the MLNode images pinned to the 0.2.14 baseline. Resolved with a pinned hardware-runtime exception: a newer upstream MLNode runtime for that host only, while the chain release and model overlay stay on the baseline (documented in the runbook release profile).
  • PoC validation timing on DAPI 0.2.14. The stock gateway could submit its first MLNode distribution in the same block as a newer store commit, missing the validation snapshot. Resolved with a ten-block PoC validation delay as an explicit test-lab genesis parameter (documented in the runbook).
  • Provider port limitations (Novita.ai). Evaluated as pay-as-you-go GPU capacity: usable for ML hosts, but does not expose the ports a full network node requires. Kept as reserve ML capacity (remaining balance to be attached to a network node's ML host later).

No hard procurement blockers were encountered during Month 1. One recurring friction: some datacenters and GPU neoclouds require KYC before activating servers, which adds delay to provisioning and slows down multi-provider deployment.

Workstream B — Testing lab (M1: Setup and QA onboarding)

M1 requirement Status
Initial test strategy Done — testing/test-strategy.md
QA search and selection Done — first engineer hired and active
Operating model Done — task board, issue templates (defect / validation request / DevNet access), SECURITY.md private disclosure path

Hiring. 60+ CVs screened, 8 candidates interviewed, one QA / infrastructure engineer hired, started on July 28. As part of onboarding he took on the DevNet deployment itself: the runbook, observability, and gateway work above; onboarding continues in Month 2. The search for the second engineer restarts in Month 2, likely with paid listing promotion; one candidate is in the pipeline and receives a test assignment once the DevNet stabilizes.

Incidents

The network spent most of the period in initial deployment and stabilization; instability during bootstrap was expected and was resolved as part of deployment work rather than tracked as incidents. Public monitoring is live (status site with per-node state and validator map, dashboards).

Spending (budget lines per proposal §14)

Line Monthly cap Spent (Month 1) Notes
DevNet machines $5,000 ≈ $1,550 ($1,065 + €420) 5 nodes + dedicated ML host + domain, itemized below
Burst GPU rental $6,500 $100 Novita.ai pay-as-you-go GPU credit (evaluated for ML hosts; balance reserved)
Tooling and reporting $250 $0
External Testing Engineers $8,000 $2,000 first QA engineer salary (paid by Aug 11)
Contingency $2,250 $0
Total $22,000 ≈ $3,650

Infrastructure breakdown (paid invoices; receipts available for review on request):

Date Provider Configuration Role Billing period Amount
07-19 DatabaseMart GPU VPS, RTX Pro 2000 16 GB node4 ML host 1 month $119.00
07-19 OneProvider dedicated CPU (EPYC/64 GB), Portland US node4 network node + public edge 1 month $77.50
07-24 DatabaseMart dedicated, RTX A5000 24 GB node0 (genesis) 1 month $254.77
07-27 OneProvider dedicated + Tesla T4 16 GB, Helsinki FI node1 1 month $331.63
07-27 Novita.ai pay-as-you-go GPU credit burst ML capacity (evaluated; balance reserved) $100.00
07-28 GIGAGPU dedicated, RTX 3090 24 GB, UK node3 1 month $270.03
07-29 NameSilo gonka-dev.net registration domain 1 year $11.85
07-29 + 08-05 LeaderGPU (LeaderTelecom) dedicated, RTX 4090 24 GB, NL node2 5 weeks €420.00

Month 2 plan (M2 targets)

  • Scale to 9+ MLNodes across target regions; publish the regional layout summary (#3, #4).
  • Publish smoke and regression checklists (#5).
  • Run the second QA engineer hiring round; we aim to close the hire, prioritizing fit over speed (#6).
  • Begin validation work as testable artifacts are handed off (validation requests).