Skip to content

[ENV-STOP] Execution host fails frozen GPU identity gate (RTX 4090 Laptop 16 GB / Windows); #56/#57/#71/#72/#73 NOT executed #74

Description

@HainingYin

Summary

This host was assigned to run the OpenLithoHub 1×RTX 4090 GPU Execution Handoff (Track A: #56 → #57; Track B: #71 → #72 → #73). On pre-flight inspection the host does not satisfy the frozen hardware gates. Per the handoff GPU identity gate (§8) and STOP rules (§14), execution STOPPED before any clone, preflight, benchmark, or source/parameter change.

Per the handoff: "GPU 不是 RTX 4090 → STOP;不要执行 #72 / #73".

All five target issues are NOT EXECUTED on this host. Nothing was fabricated; results below are directly observed hardware/OS facts, not benchmark output.

Gate-by-gate verdict

Gate Handoff requirement Observed on this host Verdict
GPU model (§4 / §8 identity gate) NVIDIA GeForce RTX 4090 (desktop, 24 GB) NVIDIA GeForce RTX 4090 Laptop GPU (16 GB) FAIL
VRAM (§4) 24 GB 15.99 GB FAIL
OS (§4) Linux Windows 11 (10.0.26200) FAIL
nvidia-smi (§16 evidence) required Failed to initialize NVML: Unknown Error FAIL
CUDA PyTorch CUDA-enabled torch 2.6.0+cu124, CUDA 12.4, cuDNN 9.1, cc 8.9 OK

The GPU is a different SKU (laptop/mobile 16 GB variant, not the desktop 24 GB), and the OS is Windows, not Linux. Both are frozen acceptance criteria, so the host is non-qualifying.

Environment evidence (directly observed, torch-reported)

  • OS: Windows-10-10.0.26200-SP0 (Windows 11), MINGW64
  • Python: 3.10.11
  • PyTorch: 2.6.0+cu124 (CUDA 12.4)
  • cuDNN: 9.1 (90100)
  • torch.cuda.is_available: True
  • device_count: 1
  • GPU name: NVIDIA GeForce RTX 4090 Laptop GPU
  • VRAM: 15.99 GB
  • SM count: 76, compute capability 8.9 (Ada)
  • nvidia-smi / nvidia-smi -q / nvidia-smi topo -m: all fail with Failed to initialize NVML: Unknown Error (NVML unavailable → driver version not readable)

Frozen commits (both verified reachable in OpenLithoHub/OpenLithoHub)

Per-issue status on this host

Issue Track Frozen commit Status Reason
#56 CUDA Host Qualification A a70d4ac NOT EXECUTED Host fails frozen GPU identity / VRAM / OS gates (§8/§14)
#57 Formal Industrial Benchmark v2 A a70d4ac NOT EXECUTED Blocked by #56 + env gate
#71 G1-4090 Scale Host Qualification B 210ed1b NOT EXECUTED Host fails GPU identity gate (not desktop RTX 4090 24 GB)
#72 G2-4090 Large-Layout Single-GPU Scale B 210ed1b NOT EXECUTED Handoff §8: "GPU 不是 RTX 4090 → 不要执行 #72/#73"
#73 G3-4090 Single-GPU Microbatch Saturation B 210ed1b NOT EXECUTED Handoff §8: "GPU 不是 RTX 4090 → 不要执行 #72/#73"

Superseded / deferred (NOT executed, per handoff §2): #66, #67, #68, #69.

What was / was not done

Done

  • Verified repo reachability, both frozen commits, and the five target issues (all present, open).
  • Captured direct hardware / OS / CUDA / PyTorch / cuDNN evidence.
  • Confirmed nvidia-smi/NVML is non-functional on this host.

Not done (by design, per STOP gate)

To proceed, a qualifying host is required

1 × NVIDIA GeForce RTX 4090   (desktop, 24 GB VRAM)
Linux
CUDA-enabled PyTorch + cuDNN
KLayout
working nvidia-smi (NVML)

Re-run from the top of the handoff on that host: #56 → #57 (a70d4ac) and #71 → #72 → #73 (210ed1b).

Cross-references

#56 #57 #71 #72 #73

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions