Runs on ≤8 GB · Hugging Face ↗
Not independently scored by Artificial Analysis — its base model AETHORIA-AI/TR-HASH-MoE-200M-160B-Refinement is the closest proxy
The smallest thing in this table by two orders of magnitude, and the name misleads: the 200M is the parameter count, the 160B is the number of training tokens. 202,731,072 parameters, 776 MiB in F32, one safetensors file. A custom architecture (TRHashForCausalLM) from a lab nobody has heard of, published 2026-08-21 under Apache-2.0, combining hash-routed mixture-of-experts with grouped-query attention.
| R9700 32GB (ROCm, transformers bf16) | 69 tok/s |
| RTX 4060 Ti 16GB (CUDA, transformers bf16) | 63 tok/s |
transformers, bf16, sdpa attention, 256 forced new tokens, median of 4 runs with the warm-up excluded. NOT comparable with the llama.cpp figures elsewhere in this table: this model has no GGUF, so it is a different runtime. One request at a time — the speed one person sees. Data-centre cards are measured but not published.
No GGUF build yet — ~0 GB estimated at 4-bit from the parameter count. Run from safetensors (transformers / vLLM / SGLang).
transformers only, and the model code comes from the repo: AutoModelForCausalLM.from_pretrained("AETHORIA-AI/TR-HASH-MoE-200M-160B-SFT", revision="e4ce55e8", trust_remote_code=True). Pin the revision rather than tracking main — the repository is four days old and has already been edited once. It brings its own chat_template.jinja. There is no llama-server path.
No GGUF exists and none can be made today, so llama.cpp, Ollama and LM Studio are all out — the architecture ships as Python in the repo. That also means running it executes the author's code on your machine (trust_remote_code), from a repository four days old with one like; we read both Python files and they are plain PyTorch, with no network, file or subprocess access, but pin the revision and run it in a container anyway. At 203M parameters do not expect it to program: we list it because it fits anywhere, not because it competes. Its English is fluent and it answers some factual questions correctly, but it also loops, confuses who is speaking, and confabulates — fine for a sentence, not for anything that depends on the content. The "sparse MoE" in the name does not apply at inference: it runs all four experts on every token and masks the ones it does not need, so it is bound by kernel launches, not by arithmetic. That is why a 32 GB R9700 beats a 4060 Ti by only 9%.