UNSW Katana使用记录
最近在研究AI Compiler,需要用到GPU,我自己倒是一张GTX1650的老古董,但是最新的nvcc已经把sm_70和sm_75的机器移除,这下只能自己去vast.ai和autodl租卡,但这两个平台各有各的问题
autodl:容器形式使用,不支持Nsight Compute获取硬件计数器
vast.ai:以VM形式创建倒是可以使用Nsight Compute,但美元记价,也不敢开太长时间
我知道学校是有GPU的,但CSE的大家都说缺卡——实际也确实是这样,UNSW Katana HPC倒是有百来张卡,而且对HDR Student开放,早些时间我去申请工单,也很顺利的批复了,今天就上手试试

机器详情
先看看官网列了什么卡
- 35 GPU nodes
- 80 x Nvidia H200 141GB (13 nodes)
- 1 x Nvidia H100 94GB (1 nodes)
- 1 x Nvidia GH200 480GB + 96GB (1 node)
- 32 x Nvidia L40S 48GB (7 nodes)
- 24 x Nvidia A100 40GB (3 nodes)
- 32 x Nvidia V100 32GB (8 nodes)
- 16 x Nvidia RTX Pro 6000 96GB (2 nodes)
- 20 GPU nodes have priority access for the school/group that purchased them
- 15 GPU nodes are for use by all researchers
好吧,百来张卡确实不算多,全UNSW分下来也就比没有强点🧐
配套的样例很详细,应该是用PBS调度的:Running Jobs on Katana
另外想用容器应该也是OK的,虽然手册没说,但确实是有权限访问apptainer的,具体使用可以看apptainer的手册:Introduction to Apptainer
(本来想跑个Nsight Compute测试和OpenAI Triton测试,但排显卡的速度确实太慢了,有空再更新,不指定显卡的情况下,应该跑的是V100) 更新已完成
存储使用
默认位置15GB,另外配上128.0GB大存储盘
Own Own Own Own Own Own Inode Inode Inode Usage Quota % Used Usage Quota % Used /home/z1234123 9.3GB 15.0GB 62% 92K -- -- /srv/scratch/z1234123 0KB 128.0GB 0% 1 500K 0%Nsight Compute(翻车)
Nsight Compute跑完了,结论是运行不了一点,vast.ai该上还是得上
Nsight Compute还不能用2023版本的,下面测试用的是2024版本的
z1234123@katana2:~/ncu-test $ cat myjob.pbs.o9316631Tue Sep 15 08:04:20 2026+-----------------------------------------------------------------------------------------+| NVIDIA-SMI 580.173.02 Driver Version: 580.173.02 CUDA Version: 13.0 |+-----------------------------------------+------------------------+----------------------+| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC || Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. || | | MIG M. ||=========================================+========================+======================|| 0 Tesla V100-SXM2-32GB On | 00000000:AF:00.0 Off | 0 || N/A 33C P0 41W / 300W | 0MiB / 32768MiB | 0% Default || | | N/A |+-----------------------------------------+------------------------+----------------------+
+-----------------------------------------------------------------------------------------+| Processes: || GPU GI CI PID Type Process name GPU Memory || ID ID Usage ||=========================================================================================|| No running processes found |+-----------------------------------------------------------------------------------------+ WARN cache for Repodata at /home/z1234123/.cache/rattler/cache/repodata is on a network/parallel filesystem (NFS/SMB/FUSE/BeeGFS/Lustre/GPFS/CephFS), redirected to /scratch/pbs.9316631.kman.restech.unsw.edu.au/pixi-cache-z1234123/repodata for this run. Set [cache.repodata] in config.toml or PIXI_CACHE_DIR to override, or [cache.netfs-redirect] = "never" to keep the original path.NVIDIA (R) Nsight Compute Command Line ProfilerCopyright (c) 2018-2024 NVIDIA CorporationVersion 2024.3.2.0 (build 34861637) (public-release) WARN cache for Repodata at /home/z1234123/.cache/rattler/cache/repodata is on a network/parallel filesystem (NFS/SMB/FUSE/BeeGFS/Lustre/GPFS/CephFS), redirected to /scratch/pbs.9316631.kman.restech.unsw.edu.au/pixi-cache-z1234123/repodata for this run. Set [cache.repodata] in config.toml or PIXI_CACHE_DIR to override, or [cache.netfs-redirect] = "never" to keep the original path.==PROF== Connected to process 1172731 (/home/z1234123/ncu-test/ncu_test)==ERROR== ERR_NVGPUCTRPERM - The user does not have permission to access NVIDIA GPU Performance Counters on the target device 0. For instructions on enabling permissions and to get more information see https://developer.nvidia.com/ERR_NVGPUCTRPERMAllocating 64.00 MB per arrayResult sample: 3.000298==PROF== Disconnected from process 1172731
================================================================================ Resource Usage on 15/09/2026 08:04:26
Job Id: 9316631Queue: CSEWalltime: 00:00:03 (requested 12:00:00)Job unsuccessful. Exit Status 1 - Exit status refers to the exit value of thetop process in the job, typically the shell.This may be the exit value of the last command executed in the shell.See job output error.
--------------------------------------------------------------------------------| | GPUs | Memory |--------------------------------------------------------------------------------| Node GPU ID | Requested Used Efficiency | Available Used || k106 3 | 1 0.04 4% | 32768 MiB 0.0B |--------------------------------------------------------------------------------|Total | 1 0.04 4.0% | |--------------------------------------------------------------------------------
--------------------------------------------------------------------------------| | CPUs | Memory |--------------------------------------------------------------------------------| Node | Requested Used Efficiency | Requested Used Efficiency || k106 | 1 0.67 67.0% | 4.0gb 0.42gb 10.5% |--------------------------------------------------------------------------------OpenAI Trtion运行
OpenAI Triton现在只能跑在sm_80以上(AI说3.3开始就移除了Volta和Turing的支持),需要手动指定显卡:
pbsnodes -av | grep gpu_modelqsub -I -l select=1:ncpus=2:mem=8gb:ngpus=1:gpu_model=A100想排到A100还是有点困难的,平均需要等6个多小时,但运行确实没问题
我用uv配置的启动环境,一个可行的pyproject.toml如下
[project]name = "triton-test"version = "0.1.0"description = "Add your description here"readme = "README.md"requires-python = ">=3.6"dependencies = [ "numpy>=1.19.5", "torch>=1.5.1", "triton>=3.1.0",]
[project.scripts]triton-test = "triton_test:main"
[build-system]requires = ["uv_build>=0.12.13,<0.13.0"]build-backend = "uv_build"启动代码(第一次使用可能还需要uv python install 3.12):
uv inituv venv --python 3.12uv sync测试用Python代码:
import torchimport tritonimport triton.language as tl
@triton.jitdef vector_add_kernel( x_ptr, y_ptr, out_ptr, n_elements: tl.constexpr, BLOCK_SIZE: tl.constexpr,): pid = tl.program_id(axis=0)
offsets = pid * BLOCK_SIZE + tl.arange(0, BLOCK_SIZE) mask = offsets < n_elements
x = tl.load(x_ptr + offsets, mask=mask) y = tl.load(y_ptr + offsets, mask=mask)
tl.store( out_ptr + offsets, x + y, mask=mask, )
def main(): print("PyTorch version :", torch.__version__) print("Triton version :", triton.__version__) print("CUDA available :", torch.cuda.is_available())
if not torch.cuda.is_available(): raise RuntimeError("CUDA is not available")
device = torch.cuda.current_device() props = torch.cuda.get_device_properties(device)
print("GPU :", props.name) print( "Compute capability:", f"{props.major}.{props.minor}", )
n = 10_000_003
x = torch.randn( n, device="cuda", dtype=torch.float32, )
y = torch.randn( n, device="cuda", dtype=torch.float32, )
output = torch.empty_like(x)
BLOCK_SIZE = 1024
grid = ( triton.cdiv(n, BLOCK_SIZE), )
vector_add_kernel[grid]( x, y, output, n, BLOCK_SIZE=BLOCK_SIZE, )
torch.cuda.synchronize()
expected = x + y
max_error = ( output - expected ).abs().max().item()
correct = torch.allclose( output, expected, rtol=1e-5, atol=1e-6, )
print("Max error :", max_error) print("Result correct :", correct)
if not correct: raise RuntimeError( "Triton kernel produced incorrect result" )
print("\nTriton works correctly on this GPU.")
if __name__ == "__main__": main()用于提交测试的PBS
#!/bin/bash
#PBS -l select=1:ncpus=1:mem=4gb:ngpus=1:gpu_model=A100#PBS -l walltime=12:00:00#PBS -M z1234123@ad.unsw.edu.au#PBS -m ae#PBS -j oe
cd $PBS_O_WORKDIR
uv run python test_triton_a100.py测试结果:
z1234123@katana2:~/triton-test $ cat triton.pbs.o9317745PyTorch version : 2.14.0+cu130Triton version : 3.8.0CUDA available : TrueGPU : NVIDIA A100-PCIE-40GBCompute capability: 8.0Max error : 0.0Result correct : True
Triton works correctly on this GPU.
================================================================================ Resource Usage on 15/09/2026 19:27:40
Job Id: 9317745Queue: CSEWalltime: 00:00:09 (requested 12:00:00)Job execution was successful. Exit Status 0.
--------------------------------------------------------------------------------| | GPUs | Memory |--------------------------------------------------------------------------------| Node GPU ID | Requested Used Efficiency | Available Used || k092 2 | 1 0.01 1% | 40960 MiB 814.0M |--------------------------------------------------------------------------------|Total | 1 0.01 1.0% | |--------------------------------------------------------------------------------
--------------------------------------------------------------------------------| | CPUs | Memory |--------------------------------------------------------------------------------| Node | Requested Used Efficiency | Requested Used Efficiency || k092 | 1 0.67 67.0% | 4.0gb 2.1gb 52.5% |--------------------------------------------------------------------------------白嫖玩玩可以,真要做Benchmark估计得等死
TileLang运行
之所以想到Tilelang,是因为Tilelang现在还支持sm_70和sm_75
uv inituv venv --python 3.12uv pip install --no-cache torch==2.14.0 --index-url https://download.pytorch.org/whl/cu126uv pip install --no-cache tilelang测试代码:
import torchimport tilelangimport tilelang.language as T
print("=" * 60)print("Environment")print("=" * 60)
print("TileLang version:", tilelang.__version__)print("PyTorch version :", torch.__version__)print("CUDA available :", torch.cuda.is_available())
if not torch.cuda.is_available(): raise RuntimeError("CUDA GPU is not visible inside PBS job")
print("GPU name :", torch.cuda.get_device_name(0))print("CUDA capability :", torch.cuda.get_device_capability(0))
major, minor = torch.cuda.get_device_capability(0)
if (major, minor) != (7, 0): print( f"WARNING: expected V100 / sm_70, " f"but got sm_{major}{minor}" )
print("=" * 60)print("TileLang vector-add test")print("=" * 60)
@tilelang.jit(target={"kind": "cuda", "arch": "sm_70"})def vector_add(N: int, block: int = 256):
@T.prim_func def kernel( A: T.Tensor((N,), "float32"), B: T.Tensor((N,), "float32"), C: T.Tensor((N,), "float32"), ): with T.Kernel(T.ceildiv(N, block), threads=block) as bx: for tx in T.Parallel(block): idx = bx * block + tx C[idx] = A[idx] + B[idx]
return kernel
N = 1 << 20
a = torch.randn(N, device="cuda", dtype=torch.float32)b = torch.randn(N, device="cuda", dtype=torch.float32)c = torch.empty_like(a)
print("Compiling TileLang kernel for sm_70 ...")
kernel = vector_add(N)
torch.cuda.synchronize()
print("Running kernel ...")
kernel(a, b, c)
torch.cuda.synchronize()
ref = a + b
torch.testing.assert_close(c, ref)
max_error = (c - ref).abs().max().item()
print("Max error:", max_error)print()print("PASS: TileLang successfully compiled and ran on V100 / sm_70.")PBS可行文件
#!/bin/bash
#PBS -l select=1:ncpus=1:mem=4gb:ngpus=1:gpu_model=V100#PBS -l walltime=01:00:00#PBS -M z1234123@ad.unsw.edu.au#PBS -m ae#PBS -j oe
cd $PBS_O_WORKDIR
module load gcc/12.2.0module load cuda/12.8
uv run python test_tilelang_v100.py测试结果:
============================================================Environment============================================================TileLang version: 0.1.14PyTorch version : 2.14.0+cu126CUDA available : TrueGPU name : Tesla V100-SXM2-32GBCUDA capability : (7, 0)============================================================TileLang vector-add test============================================================Compiling TileLang kernel for sm_70 ...2026-09-15 22:10:27 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:137): TileLang begins to compile kernel `kernel` with `out_idx=None`2026-09-15 22:10:29 [TileLang:tilelang.jit.kernel:INFO] (kernel.py:145): TileLang completes to compile kernel `kernel`Running kernel ...Max error: 0.0
PASS: TileLang successfully compiled and ran on V100 / sm_70.
================================================================================ Resource Usage on 15/09/2026 22:10:46
Job Id: 9319953Queue: CSEWalltime: 00:00:26 (requested 01:00:00)Job execution was successful. Exit Status 0.
--------------------------------------------------------------------------------| | GPUs | Memory |--------------------------------------------------------------------------------| Node GPU ID | Requested Used Efficiency | Available Used || k105 1 | 1 0.0 0% | 32768 MiB 410.0M |--------------------------------------------------------------------------------|Total | 1 0.0 0.0% | |--------------------------------------------------------------------------------
--------------------------------------------------------------------------------| | CPUs | Memory |--------------------------------------------------------------------------------| Node | Requested Used Efficiency | Requested Used Efficiency || k105 | 1 0.31 31.0% | 4.0gb 0.7gb 17.5% |--------------------------------------------------------------------------------uv使用注意事项
uv使用的cache可能会把15GB的小空间挤爆,建议还是直接使用conda或python venv解决(或是uv开启--no-cache)
Tips
qstat -u "$USER"
可以查看已经提交的所有任务
module avail gcc
查询可用的GCC,Tilelang似乎使用CuTe模板,GCC版本小于10会无法运行
结语
原本计划还要去试试CSE的HPC的,但CSE存储每人共享上课用的3GB存储容量,随便上个Conda环境就没了——只能期待他们后续给更多Storage了