Phi-3-mini on raw WebGPU
Runs in your browser on hand-written WGSL kernels. Weights download once and cache locally; prompts and generation stay on this machine.
~2 GB
weights
~70 t/s
M2 Max
context
Pick a different model →
Not now
Preparing Zero-TVM
Starting…
0%
Details

      
hand-written WGSL · paged KV · WebGPU
Waiting
Chat with Phi-3-mini
q4f16_1 · 3.8 B 4 K context Runs on your GPU · WebGPU On-device only
Running on your own GPU through hand-written WGSL kernels. Prompts, tokens and the KV cache stay in this tab.
Enter to send · Shift+Enter for new line · Zero TVM · 10 WGSL kernel roles