Popular repositories Loading
-
-
-
KernelBench-v2
KernelBench-v2 PublicForked from ScalingIntelligence/KernelBench
KernelBench v2: Can LLMs Write GPU Kernels? - Benchmark with Torch -> Triton (and more!) problems
-
ai-scientist-artefacts-v1
ai-scientist-artefacts-v1 PublicArtefacts from the first complete run of the Lossfunk AI Scientist pipeline for paper accepted at Agents4Science 2025.
-
Repositories
- crafter-symbolic-solver Public
A pure-Python heuristic agent for Crafter, exploring diamond collection and achievement scores within 10,000 steps. Reads RGB observations, keeps memory and runs on CPU. Includes benchmarks, a live viewer and GIF export.
- verbalizing-eval-awareness Public
A fun experiment to see whether models are self-aware and can actively state that they are being evaluated or not.
- IR-vOICe Public
Sensory substitution: IR camera → vOICe soundscape. Let a user see IR scenes through sound.
- vlm_benchmark Public
- research-taste-eval-v1 Public
Data and analysis code for 'Eliciting Research Taste in LLMs through Future Research Direction Choice' (COLM 2026, LM4SCI Workshop).
- Denoising-models-illusion-representations Public
This is the supporting code and data for the Paper "Denoising Models Develop Human-Like Perceptual Illusion Representations Across Architectures"
- esolang-metaprogramming Public
Code for 'Frontier Coding Agents Use Metaprogramming to Adapt to Unfamiliar Programming Languages' (Lossfunk, 2026) — EsoLang-Bench harness, interpreters, and reproducible experiments. Run with Claude Code, Codex, OpenCode, or any model via OpenRouter.
- Hybrid-Neural-World-Models Public
- EsolangBench Public
Top languages
Loading…
Most used topics
Loading…