Software Engineer and maker of many things.
-
Build Things
- Western Australia
- https://www.jayleaton.com/
- @build-things
- @jayleaton
Highlights
- Pro
Pinned Loading
-
-
glm53-tensorfold-spark
glm53-tensorfold-spark PublicGLM-5.3-Flash (abliterated EXL3) on 2x NVIDIA DGX Spark with the TensorFold engine: 1.8x faster decode than vLLM, 4x256k concurrent threads, byte-exact speculative decoding. Work in progress.
-
qwen38-flash-next-exl3-spark
qwen38-flash-next-exl3-spark PublicServe Qwen3.8-Flash-Next (uncensored EXL3) on one NVIDIA DGX Spark: one Docker image, OpenAI-compatible endpoint, MTP speculative decoding, 262k context. ~80 tok/s.
Python 1
-
qwen-image21-tensorfold-rtx
qwen-image21-tensorfold-rtx PublicQwen-Image 2.1 on TensorFold's NVFP4 kernels: a drop-in ComfyUI loader, 4x faster than the GGUF path on one RTX 5070 Ti (native Windows). Work in progress.
Python 7
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.




