Skip to content

Sandboxed flash attention benchmarks #2614

Description

@florianscheidl

Describe the task. It can be a feature, documentation, etc.

We want to assess the impact of different flash attention versions on throughput and memory consumption for different token size configurations.
The hope is for us to be able to do this in a sandboxed script, shielding much of WeatherGenerator's complexity. At the same time, our tests must reflect on the token configurations distributions in WeatherGenerator models or, if not, this limitation must be kept in mind.

Sandboxed environments in which we abstract away parts of the WG codebase would also be helpful for other performance-related investigations, e.g., testing other model sharding strategies.

Hedgedoc URL, if you are keeping notes, plots, logs in hedgedoc.

No response

Area

  • datasets, data readers, data preparation and transfer
  • model
  • science
  • infrastructure and engineering
  • evaluation, export and visualization
  • documentation
  • performance

Metadata

Metadata

Assignees

No one assigned

    Labels

    performanceWork related to performance improvements

    Type

    No type

    Fields

    No fields configured for issues without a type.

    Projects

    Status
    No status

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions