{"v":1,"id":"model:hf:oruk/orukeet","slug":"model-oruk-orukeet","kind":"model","category":"speech","title":"orukeet","summary":"Nathan Roll1,2 · Irene Yi1,2 · Büşra Marşan1,2 Vianney Grenez1 · Gabriel Stein4 · Momcilo Mrkaic5 Pavle Padjin5 · Vladimir Zeljkovic5 · Calbert Graham1,3","source":{"provider":"hf","ref":"oruk/orukeet","url":"https://huggingface.co/oruk/orukeet","rev":"b59c13a733fce5cf193230d0fa317ee7f145a108","fetchedAt":"2026-10-09T19:01:06.454Z","etag":"W/\"6e6d-8Pkwzm+B787qWCqYFrShFXh2axw\""},"author":{"name":"oruk","url":"https://huggingface.co/oruk"},"license":{"spdx":"cc-by-sa-4.0","raw":"cc-by-sa-4.0","open":true},"metrics":{"downloads":40095,"downloadsWeek":15664,"likes":99,"stars":54249,"openIssues":352,"lastRelease":{"tag":"v1.9.5","at":"2026-10-06T17:39:17Z"},"pushedAt":"2026-10-06T17:39:15Z","takenAt":"2026-10-09T19:01:06.454Z"},"tags":["nemo","onnx","safetensors","gguf","parakeet_tdt","parakeet","tdt","sherpa-onnx","multilingual","speech-recognition","gabor","fastconformer","automatic-speech-recognition","bg","hr","cs","da","nl","en","et","fi","fr","de","el","hu","it","lv","lt","mt","pl","pt","ro","ru","sk","sl","es","sv","uk"],"pipeline":"automatic-speech-recognition","links":{"github":"ggml-org/whisper.cpp","npm":"nodejs-whisper"},"updatedAt":"2026-09-24T20:33:55.000Z","collectedAt":"2026-10-09T19:01:06.454Z","review":{"numbers":["40,095 downloads on Hugging Face","99 likes","license cc-by-sa-4.0","0.7 GB for orukeet-transcribe-cpp-Q8_0.gguf","54,249 stars on ggml-org/whisper.cpp","352 open issues and PRs","last release v1.9.5 on 2026-10-06","15,664 npm downloads a week for nodejs-whisper","latest nodejs-whisper@0.3.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-09T19:01:06.454Z","http":200},"description":"# Orukeet\n\nNathan Roll1,2 · Irene Yi1,2 · Büşra Marşan1,2\nVianney Grenez1 · Gabriel Stein4 · Momcilo Mrkaic5\nPavle Padjin5 · Vladimir Zeljkovic5 · Calbert Graham1,3\n\n1 Oruk AI\n\n2 Stanford University\n3 University of Cambridge\n4 OpenWhispr\n5 Hoid\n\nOrukeet is a 25-language speech recognizer built from NVIDIA Parakeet TDT 0.6B v3. It replaces half of the encoder's temporal depthwise filters with **12,288 fitted, frozen Gabor kernels** and trains the remaining parameters on multilingual and multi-accent data.\n\nOrukeet outperforms Parakeet on **61 of 74 tested splits**, including LibriSpeech test-clean (**1.46% vs. 1.53% WER**), test-other (**2.86% vs. 3.14%**), and FLEURS English (**3.82% vs. 4.28%**). Across all 25 FLEURS languages, pooled WER is **9.85% vs. 11.01%**, a **10.6% relative reduction**. Final adaptation and checkpoint selection use LibriSpeech test-other.\n\nUse Orukeet for recordings, media, batch transcription, server workers and interactive applications. NeMo, ONNX INT8, native Q8 and native F16 all derive from the same **r3 release checkpoint** (`031c8ddab484`).\n\n[Code](https://github.com/Oruk-AI/orukeet) · [OpenWhispr PR](https://github.com/OpenWhispr/openwhispr/pull/2085) · [Technical report](orukeet-technical-report.pdf) · [Artifact hashes](ARTIFACTS.json)\n\n## Run Orukeet with Transformers\n\nA standard FP32 Transformers export is available at the repository root. It uses\n`ParakeetForTDT` without custom remote code and works with Buzz's existing Hugging\nFace model option. See [setup, conversion provenance and runtime qualification](transformers/README.md).\nThe NeMo evaluation below remains the source-model benchmark; the Transformers\nexport has separate compatibility measurements.\n\n## Run Orukeet with NeMo\n\nUse a CUDA-enabled PyTorch environment with `nemo_toolkit[asr]==3.0.0` and `huggingface-hub`. The [recorded source environment](https://github.com/Oruk-AI/orukeet/blob/main/evidence/standard-asr-20260908/r…\n\nSource: https://huggingface.co/oruk/orukeet","install":{"kind":"model","hfId":"oruk/orukeet","gated":false,"format":"gguf","files":[{"name":"model.safetensors","size":2508311120,"sha256":"092cff73445117d4d234c5e4fb37d86ad16a348be3426e10ff38c79d35758569"},{"name":"onnx/combined-v0.1.0-int8/decoder_joint-model.int8.onnx","size":18202844,"sha256":"95d3b1f53f9aadc5ef58e63664a3681a2184ee228b5010e1ef975a1c4ea8318a"},{"name":"onnx/combined-v0.1.0-int8/encoder-model.int8.onnx","size":653182378,"sha256":"7b55f2a504a20a8e462899f5befd45f4a1784948d76ed0127902d9cf39405487"},{"name":"onnx/combined-v0.1.0-int8/nemo128.onnx","size":139764,"sha256":"a9fde1486ebfcc08f328d75ad4610c67835fea58c73ba57e3209a6f6cf019e9f"},{"name":"onnx/sherpa-v0.1.0-int8/decoder.int8.onnx","size":11845332,"sha256":"c185c2afb4c77c94bb1314807ecb3dc1623057a3dc540b83e10301af9bf4cfca"},{"name":"onnx/sherpa-v0.1.0-int8/encoder.int8.onnx","size":653182378,"sha256":"7b55f2a504a20a8e462899f5befd45f4a1784948d76ed0127902d9cf39405487"},{"name":"onnx/sherpa-v0.1.0-int8/joiner.int8.onnx","size":6355335,"sha256":"1a7e90abf7172d926dd7e2edac2a5d5035c24dfb641a15e131d57b6a5f63cdd3"},{"name":"orukeet-transcribe-cpp-Q8_0.gguf","size":739508608,"quant":"Q8_0","sha256":"cad2f52ac91cad829279422301989687c2cf02e19157352ed25ea501b90dbb7e"},{"name":"orukeet-v0.1.0-f16.gguf","size":1296681088,"quant":"F16","sha256":"de53fb8ec251fb07ade15baabe17b00774ae3f1112f8618b062337f90fb49194"},{"name":"orukeet-v0.1.0-q8.gguf","size":714456704,"sha256":"93ce19c6d8244acbfea980eeaf970531d4f216171578ef8e041dcc2d070a45bd"},{"name":"transcribe-cpp/orukeet-Q8_0.gguf","size":739508608,"quant":"Q8_0","sha256":"cad2f52ac91cad829279422301989687c2cf02e19157352ed25ea501b90dbb7e"}],"totalBytes":7341374159,"suggestedFile":"orukeet-transcribe-cpp-Q8_0.gguf","requirements":{"ramGb":2,"diskBytes":739508608,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["whisper.cpp"],"command":"lsh models install hf:oruk/orukeet"}}