Skip to content

Streamed Multi-Head Attention for Altera Backend - #1551

Draft
bugracyln wants to merge 14 commits into
fastmachinelearning:mainfrom
bugracyln:streamed_MHA
Draft

bugracyln wants to merge 14 commits into
fastmachinelearning:mainfrom
bugracyln:streamed_MHA

Conversation

@bugracyln

Copy link
Copy Markdown

Description

📝 Please include a summary of the change.

Type of change

For a new feature or function, please create an issue first to discuss it
with us before submitting a pull request.

Note: Please delete options that are not relevant.

  • Bug fix (non-breaking change that fixes an issue)
  • Documentation update
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • A new research paper code implementation
  • Other (Specify)

Tests

📝 Please describe the tests that you ran to verify your changes.

  • The module has been tested extensively with various test scripts, both for accuracy testing by using the bridge to use .predict() feature to compare with the actual front-end-produced multi-head attention layer. This module has also been tested on an actual FPGA inside a small LLM architecture and was confirmed to provide correct results.

Checklist

  • I have read the guidelines for contributing.
  • I have commented my code, particularly in hard-to-understand areas.
  • I have made corresponding changes to the documentation.
  • My changes generate no new warnings.
  • I have installed and run pre-commit on the files I edited or added.
  • I have added tests that prove my fix is effective or that my feature works.

@bugracyln
bugracyln marked this pull request as draft September 17, 2026 02:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants