GitHubResearch & data
GitHub
375 builds · page 2 of 10
Auditable Qwen3.5-0.8B pipeline covering data synthesis, training, calibration, fixed evaluation, Jev-compatible serving, and public model and dataset artifacts; its metrics are…
GitHubResearch & data
JevForge
GitHubResearch & data
jevcal
Local decision cache keyed on (model, schema, state), with redaction and canonicalisation before hashing, for cheaper repeats and deterministic replay in CI.
GitHubResearch & data
jevcache
Chinese/English retrieval evaluation comparing Jev reranking with lexical, embedding, and fusion baselines, including judge-circularity analysis.
GitHubResearch & data
jev-search-rerank-eval
Retrieval and RAG: uses Jev Noul judgments to assess retrieved documents for relevance and usefulness as answer evidence, then sorts results and optionally filters them using a…
GitHubResearch & data
jev-reranker
Fourteen-dataset reranking study with saved raw responses, paired bootstrap intervals, order-sensitivity checks, and no-relevant-document tests; comparisons remain study-specific.
GitHubResearch & data
jev-rerank-bench
Reproducible experiments on whether reranking with Jev improves a small RAG system, on a locked Turkish dataset, with quality, latency and cost reported together.
GitHubResearch & data
jev-rag-benchmark
Measures whether ORDER BY over a Jev probability is defensible (inversion rate, Score ordinality against a human grade, calibration, wording invariants, sort-key ties) under a…
GitHubResearch & data
jev-orderby-bench
Small Korean/English sample study with recorded responses, including medical-text questions; not a clinical validation.
GitHubResearch & data
jev-korean-benchmark
Independent Jev versus GPT-5.6 Terra comparison on three labeled classification tasks, reporting accuracy, calibration, latency, and cost.
GitHubResearch & data
jev-eval
GitHubResearch & data
jev-benchmarks
Independent synthetic-task study of Jev 1.13.0 framing sensitivity and failures, with raw responses and offline report checks.
GitHubResearch & data
jev-behavior-study
Compares a typed decision model with classical classification pipelines across eight datasets, with a published protocol and an interactive report.
GitHubResearch & data
Jev vs. ML
Educational MLX/Qwen vision-language experiment sharing image context across candidate-scoring questions; its probabilities are not calibrated correctness estimates.
GitHubResearch & data
Jev Visual
Self-hosted implementation of Jev's System One API on the 400M-parameter GLiFormer model.
GitHubResearch & data
jeff
GitHubResearch & data
Janus
BAML language support for an AI if-statement: `.feels()` as a real, typed method backed by a decision model.
GitHubResearch & data
feelings
Scores a finite set of choices with a model you already run in llama.cpp and returns a typed decision with a probability distribution.
GitHubResearch & data
choosekit
Open data, training recipe, and a 9B model for Jev-style choice and true/false decisions on Apple Silicon or NVIDIA GPUs.
GitHubResearch & data
Bespoke Nimble
GitHubGames & real time
typesafe-snake
GitHubGames & real time
typesafe-mario
GitHubGames & real time
tsai-sc
Jev `Choice` judgments answer lateral-thinking puzzle questions and assess proposed solutions, while application code requires supported facts, a coherent explanation, and…
GitHubGames & real time
Soupbase
GitHubGames & real time
killmyidea
Town of 10,000 computed personas where Jev scores a post, listing, product, or headline against about 60 audience attributes to plan who sees it, then answers one batched Choice…
GitHubGames & real time
Jevtown
RuneBench-based RuneScape harness that maps Jev choices to a bounded game-action catalog and records tick-level results.
GitHubGames & real time
JevScape
Three.js driving simulation where Jev chooses among candidate paths and speeds while local code handles vehicle dynamics and geometry.
GitHubGames & real time
JevPilot
Tetris where deterministic code enumerates every reachable placement, including tucks and spins, and writes each one as an English sentence; Jev returns a probability for all of…
GitHubGames & real time
jev-tetris
Pokemon Red on PyBoy where deterministic code owns the route and arithmetic and Jev picks only at branches, with every battle turn's faint prediction scored by Brier against the…
GitHubGames & real time
jev-plays-pokemon-red
Command-palette demo where a single Choice question over a 77-command catalog turns the returned probability distribution into the per-keystroke ranking for Portuguese or English…
GitHubGames & real time
jev-palette
Fine-grained robot control on LIBERO tasks with physics previews and configurable task definitions.
GitHubGames & real time
jev-libero
Collection of inspectable Jev demos, including scripted support conversations with typed intent, escalation, and suggested-response decisions.
GitHubGames & real time
jev-experiments
Voice and finger-pointing control of a tldraw canvas: Jev picks the action, target shape and place from each partial transcript plus the fingertip position; deterministic code…
GitHubGames & real time
jev-canvas
A 2D top-down car in the browser that turns its sensors into a JSON state every 200 ms and executes four typed answers.
GitHubGames & real time
Jev Self-Driving Sim
Multi-drone simulation where typed reflexes fly the fleet and an optional slower planner may advise but never takes control.
GitHubGames & real time
Jev Reflex Autonomy Lab
Recorded chess experiments with a candid result: Jev on its own still blunders pieces.
GitHubGames & real time
Jev Chess Lab
Browser stealth game where Jev judges guards while deterministic code owns the world.
GitHubGames & real time
heist-one
Serves DiffusionGemma-Jev behind a compatible API on Cloud Run, with a small game demo on top.
GitHubGames & real time
djev-run
Three.js quickscope arena where Jev decides movement, aiming, ADS, firing, and jumping at roughly 9 Hz.
GitHubGames & real time