Skip to content

Enable prefetch iteration - #382

Closed
t4c1 wants to merge 35 commits into
intel:mainfrom
t4c1:npot_prefetch
Closed

t4c1 wants to merge 35 commits into
intel:mainfrom
t4c1:npot_prefetch

Conversation

@t4c1

@t4c1 t4c1 commented May 19, 2025 •

Copy link
Copy Markdown

Enables iteration of prefetch atom to cover prefetch tile. In other words relaxes the requirement for the prefetch tile size to match prefetch atom size.

This is done by using the same path for prefetch that nvidia code uses - going through copy implementation.

This PR also:

  • includes a lot of bugfixes for prefetch atom layouts
  • adds some missing copy traits
  • removes some duplicated prefetch implementations (_V atoms duplicating what is already in _N)

Comment thread include/cute/atom/copy_traits.hpp Outdated
Comment thread include/cute/atom/copy_traits_xe.hpp Outdated
Comment thread include/cute/atom/copy_traits_xe.hpp Outdated
Comment thread include/cute/arch/xe_copy_1B.hpp Outdated
Comment thread include/cute/atom/copy_traits_xe.hpp Outdated

@joeatodd joeatodd left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM 👍

@rolandschulz

Copy link
Copy Markdown

There is no need to finish this giving the move to the rearchitected atoms.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants