Jevstiller: distill Jev calls into a local model
A self-hosted proxy that trains a local model on Jev's answers. It served 72% of Banking77 locally at 15 ms, with a 2% audit and a bound on disagreement.
Jev, distilled on the fly.</b> Same call. Same answers. Your hardware.</p> <p align="center"> <a href="https://pypi.org/project/jevstiller/"><img alt="PyPI" src="https://img.shields.io/pypi/v/jevstiller?color=137572&label=pypi"></a> <a href="https://pypi.org/project/jevstiller/"><img alt="Python 3.10+" src="https://img.shields.io/pypi/pyversions/jevstiller?color=137572"></a> <a href="https://github.com/tomerglick57/Jevstiller/actions/workflows/ci.yml"><img alt="CI" src="https://img.shields.io/github/actions/workflow/status/tomerglick57/Jevstiller/ci.yml?branch=main&label=ci"></a> <a href="https://github.com/tomerglick57/Jevstiller/pkgs/container/jevstiller"><img alt="Container image" src="https://img.shields.io/badge/ghcr.io-jevstiller-137572"></a> <a href="LICENSE"><img alt="License" src="https://img.shields.io/github/license/tomerglick57/Jevstiller?color=137572"></a> <a href="https://jevstiller.pages.dev"><img alt="Docs" src="https://img.shields.io/badge/docs-jevstiller.pages.dev-137572"></a> </p> Put Jevstiller in front of a repeated [Jev](https://docs.typesafe.ai) classification call. At first every request still goes to Jev. From Jev's own answers — with their full probability distributions — it trains a small local model on your traffic, checks that the model agrees with Jev within a budget you set, and then answers most requests itself. Uncertain or novel input, and a permanent random audit slice, keep going to Jev. <picture> <source media="(prefers-color-scheme: dark)" srcset="https://github.com/tomerglick57/Jevstiller/raw/main/docs/media/flow-dark.svg"> <img alt="Your services call Jevstiller instead of Jev. Most requests are answered by the local model in about 16 ms; the uncertain ones and a 2% audit go on to Jev." src="https://github.com/t