百科.dev
全部条目AI 编程趋势榜开源项目技术资讯提交条目
登录
返回工具页/返回 Issues 列表
#131·FlexLLMGen

如何将分析结果与成本模型的参数相匹配?

作者: xvanQ创建于 2024年1月31日更新于 2024年4月30日

配置文件带宽的输出如下: size: 0.25 MB, gpu-to-cpu bandwidth: 5.505 GB/s size: 32.00 MB, gpu-to-cpu bandwidth: 13.220 GB/s size: 128.00 MB, gpu-to-cpu bandwidth: 13.324 GB/s size: 0.25 MB, cpu-to-gpu bandwidth: 4.556 GB/s size: 32.00 MB, cpu-to-gpu bandwidth: 12.285 GB/s size: 128.00 MB, cpu-to-gpu bandwidth: 12.251 GB/s 其中 ctog_bdw 是 ctog_bdw_cache 是 gtoc_bdw_hidden

内容来源: FMInference/FlexLLMGen

查看 GitHub 原文在 GitHub 查看讨论