* ggml-cuda: fix divergent barrier in f16 flash attention * ggml-cuda: avoid duplicate metadata pointer setup