nvtx
Here are 14 public repositories matching this topic...
Hooked CUDA-related dynamic libraries by using automated code generation tools.
-
Updated
Dec 12, 2023 - C
A safe Rust FFI binding for the NVIDIA® Tools Extension SDK (NVTX).
-
Updated
Jan 18, 2024 - Rust
PyProf2: PyTorch Profiling tool
-
Updated
Jun 25, 2020 - Python
Thin pybind11 wrapper for NVTX wrappers -- with some bells and whistles attached.
-
Updated
May 6, 2021 - Python
Header-only C++17 latency profiler for real-time CUDA loops. Per-stage distributions, deadline tracking, and Perfetto timelines
-
Updated
Aug 26, 2026 - C++
🎬 Explore GPU training efficiency with FP32 vs FP16 in this modular lab, utilizing Tensor Core acceleration for deep learning insights.
-
Updated
Feb 20, 2026 - Python
Reproducible GPT-2 distributed-training benchmarks on 1-8 V100 GPUs using Slurm, PyTorch, DeepSpeed, NCCL, NVTX, and Nsight Systems.
-
Updated
Jun 25, 2026 - Python
A reproducible GPU benchmarking lab that compares FP16 vs FP32 training on MNIST using PyTorch, CuPy, and Nsight profiling tools. This project blends performance engineering with cinematic storytelling—featuring NVTX-tagged training loops, fused CuPy kernels, and a profiler-driven README that narrates the GPU’s inner workings frame by frame.
-
Updated
Apr 25, 2026 - Python
Profile-first ML systems project optimizing a multi-camera end-to-end driving model for hardware efficiency using PyTorch, CUDA streams, NVTX instrumentation, and Nsight Systems.
-
Updated
Feb 12, 2026 - Python
Profiling with Precision. Documenting with Style.
-
Updated
Apr 25, 2026 - Jupyter Notebook
Add this topic to your repo
To associate your repository with the nvtx topic, visit your repo's landing page and select "manage topics."