Change the repository type filter
All
Repositories list
25 repositories
local-kimi
PublicRun Kimi locally and use it with any coding agent. Translates between Anthropic Messages, OpenAI Chat Completions and OpenAI Responses, detected per request.runinfra-cli
Public- Faster attention kernels for serving TML's Inkling model on vLLM. 2.7x over the shipping path on H100, and the only implementation that runs on A100.
runinfra-sdk
PublicOfficial RunInfra SDK (TypeScript + Python) | optimized inference deploymentssglang
PublicMemoir
PublicWeights that never freeze. A continual-learning architecture, and the pre-registered experiment that falsified its central claim. https://arxiv.org/abs/2607.207…autotree
PublicTree execution engine for LLM inference: fork, merge, prune KV cache at token granularitybonsai-turbo
PublicSingle-launch batch-1 decode engine for PrismML Bonsai 27B (ternary and 1-bit) on NVIDIA GPUs. 1.76x the vendor llama.cpp fork on H100, same outputs.auto
Publicthe agi compiler: records llm agent behavior, proves what repeats, and compiles it into verified, sandboxed wasm binaries that run for microdollars. nothing fig…openfang
PublicOpen-source Agent Operating SystemAutoMegaKernel
PublicAn agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch-1 LLM decode, paper: h…StreamIndex
PublicMemory-bounded compressed sparse attention via streaming top-k. Triton kernels for the DeepSeek-V4 lightning indexer. 32x regime extension on a single H200 | by…ouroboros
PublicDynamic weight generation for recursive transformers via input-conditioned LoRA modulationhclsm
PublicHierarchical Causal Latent State Machines for Object-Centric World Modelingautokernel
PublicTIDE
PublicDynamic per-token early exit for LLM inference. Skip layers tokens don't needqwen3.5-triton
Publicpicolm
Publictiny-tpu
PublicMinimal TPU implementation with 8x8 systolic array and PyTorch integrationgpuci
PublicGPU CI/CD tool that tests CUDA kernels across multiple GPUs in parallel - Part of RightNowRightNow-GPU-Database
PublicRightNow-Tile
PublicOpen-source transpiler for CUDA Tile (13.1) migrationrightnow-cli
Publicgpu-profiler
Public- RightNow Arabic LLM Corpus - One of the largest high-quality Arabic text datasets for LLM training
ProTip! When viewing an organization's repositories, you can use the
props. filter to filter by custom property.