Popular repositories Loading
-
-
riscv-pipelined-core
riscv-pipelined-core Public5-stage pipelined RV32I-subset core in SystemVerilog, running on a Nexys A7. 64-entry 2-bit branch predictor, 2-way set-associative D-cache, UART echo and VGA pattern generator over MMIO.
SystemVerilog
-
cuda-tiled-sgemm
cuda-tiled-sgemm PublicSingle-precision dense matrix multiplication written from scratch in CUDA, reaching 80% of cuBLAS on an RTX 4090 through 128x128 block tiles, 8x8 register tiles, and double-buffered shared memory.
Cuda
-
flashattention-cuda
flashattention-cuda PublicFlashAttention forward pass written from scratch in CUDA: tiled keys and values, online softmax, and no materialized T x T attention matrix.
Cuda
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.