add backward for linear layer kernel, see hazyresearch implementation: https://github.com/HazyResearch/flash-attention/blob/v0.2.2/flash_attn/ops/triton/linear.py#L285
add backward for linear layer kernel, see hazyresearch implementation:
https://github.com/HazyResearch/flash-attention/blob/v0.2.2/flash_attn/ops/triton/linear.py#L285