Every retained configuration and result

Complete catalogue.

98% agent-driven model benchmarks and optimizations.

Benchmarks

Select any result row to show the configuration's applied optimizations, parameters, placement, notes, and evidence.

Needle 2

Size
45M
Quant
CQ2-bit
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
Cactus NeedleCPUTODOTODO?NANA
Cactus NeedleCPUTODOTODO?NANA
Cactus NeedleCPUTODOTODO?NANA

TinyLlama OpenOrca 1.1B

Size
1B
Quant
Q4_K_M
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cppCPU + P401024251017.3164.3NANA
llama.cpp/poolsideV10040961284150.9384.1NA29%
llama.cppV1001024252547.6201.3NA15%
llama.cppV100 + P40102425176.112.6NA0%

Bonsai

Size
8B
Quant
Q1_0
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cppCPUTODO12816.16.7NANA
llama.cppP40TODO128717.263.5NA21%
llama.cppV100TODO1281338.8123.3NA16%

Gemma 4

Size
25B / 4B
Quant
UD-Q4_K_XL
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cppCPU40962010.4NANA
llama.cppCPUTODO12831.19.6NANA
llama.cppCPU4096TODO8.8NANA
llama.cppCPU40962020.89.6NANA
llama.cppCPU409620Local CPU MTP+3.9%57.5%10.8NANA
llama.cppCPU4096TODOLocal CPU MTP+12.9%57.5%9.9NANA
llama.cppCPU409620Remote CPU full-model draft-25.2%100.0%20.47.1NANA
llama.cpp/lagunaV100TODO128776.581.0NANA
llama.cpp/poolsideV100409620336.377.4NANA
llama.cpp/poolsideCPU + V100409620Local CPU MTP-32.3%61.5%330.352.4NANA

Qwen3.8

Size
27B
Quant
Q4_K_M
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cppV100(2)40961193.132.7NA62%
llama.cppV100(2)4096512873.025.238%48%
llama.cppV100(2)4096111122.334.1NA65%
llama.cppV1004096343459.533.9NA65%
llama.cppV100(2)4096111123.633.9NA64%
llama.cppV1004096343486.032.5NA62%
llama.cppV100(2)4096111130.333.8NA64%
llama.cppV1004096343459.833.7NA64%
llama.cppV100(2)4096111121.633.9NA64%
llama.cppV1004096343487.433.9NA64%
llama.cppV100(2)4096111176.033.0NA63%
llama.cppV10040963434198.033.0NA63%
llama.cppV100(2)4096111127.233.6NA64%
llama.cppV100(2)4096111120.733.2NA63%
llama.cppV100(2)4096111124.934.2NA65%
llama.cppV100(2)4096111133.933.6NA64%
llama.cppV100(2)4096202049.933.9NA65%
llama.cppV100(2)4096242446.833.8NA64%
llama.cppV1004096343468.133.0NA63%
llama.cppV1004096343484.233.4NA63%
llama.cppV1004096343465.934.1NA65%
llama.cppV100(2)4096343494.033.5NA64%
llama.cppV100(2)4096111133.2NA63%
llama.cppV1004096343433.3NA63%
llama.cppV100(2)4096111192.632.7NA62%
llama.cppV100(2)4096111193.333.0NA63%
llama.cppV100(2)40963434239.932.9NA63%
llama.cppV100(2)40963434241.433.1NA63%
llama.cppV100(2)4096111175.933.1NA63%
llama.cppV10040963434202.132.9NA62%
llama.cppV100(2)4096111124.731.9NA30%
llama.cppV100(2)4096343475.530.5NA29%
llama.cppV100(2)4096111128.134.2NA65%
llama.cppV100(2)4096343487.234.2NA65%
llama.cppV100(2)409603090630906589.027.926%60%
llama.cppV100(2)409603090630906603.228.026%61%
llama.cppV100409603090630906687.128.030%61%
llama.cppV100409603090630906743.728.133%61%
llama.cppV100409603090630906778.128.134%61%
llama.cppV100409603090630906595.528.126%61%
llama.cppV100327683090630906591.127.826%60%
llama.cppV100409611Local GPU MTP+3.6%68.9%68.633.9NANA
llama.cppV100131072TODOLocal GPU MTPNANA
llama.cppV100131072111781111781Local GPU MTP+2.5%425.225.419%NA
llama.cppV100131072112941112941Local GPU MTP+32.4%249.432.911%NA
llama.cppV100131072116943116943Local GPU MTP+11.8%420.227.818%NA
llama.cppV100131072118103118103Local GPU MTP-2.4%243.124.211%NA
llama.cppV10040961111Local GPU MTP+42.7%74.7%26.848.7NANA
llama.cppV10040963434Local GPU MTP+43.3%74.7%54.248.9NANA
llama.cppV10040961111Local GPU MTP+43.0%74.7%20.148.4NANA
llama.cppV100(2)40963434Local GPU MTP+42.6%74.7%54.848.3NANA
llama.cppV10040961111Local GPU MTP+58.4%73.1%18.053.5NANA
llama.cppV100(2)40963434Local GPU MTP+34.3%73.4%50.845.4NANA
llama.cppV10040961111Local GPU MTP+44.0%73.1%22.948.8NANA
llama.cppV100(2)40963434Local GPU MTP+28.7%73.4%62.943.6NANA
llama.cppV10040961111Local GPU MTP+43.7%74.7%24.148.7NANA
llama.cppV10040963434Local GPU MTP+40.2%74.7%74.747.5NANA
llama.cppV10040961111Local GPU MTP+44.0%74.7%25.848.7NANA
llama.cppV10040963434Local GPU MTP+43.7%74.7%74.848.6NANA
llama.cppV10040961111Local GPU MTP+43.1%74.7%27.148.5NANA
llama.cppV10040963434Local GPU MTP+43.6%74.7%77.248.7NANA
llama.cppV10040961111Local GPU MTP+30.3%74.7%68.543.0NANA
llama.cppV10040963434Local GPU MTP+29.7%74.7%171.242.7NANA
llama.cppV10040961111Local GPU MTP+42.6%74.7%20.247.9NANA
llama.cppV10040961111Local GPU MTP+48.0%81.3%24.749.7NANA
llama.cppV10040961111Local GPU MTP+38.7%68.9%21.146.6NANA
llama.cppV10040961111Local GPU MTP+43.1%74.7%23.348.1NANA
llama.cppV10040962020Local GPU MTP+63.7%100.0%36.955.0NANA
llama.cppV10040962424Local GPU MTP+54.0%81.3%49.151.8NANA
llama.cppV10040963434Local GPU MTP+43.1%74.7%69.748.1NANA
llama.cppV10040963434Local GPU MTP+50.9%82.7%73.550.7NANA
llama.cppV10040963434Local GPU MTP+44.8%78.6%69.248.7NANA
llama.cppV100(2)40963434Local GPU MTP+41.9%74.7%59.247.7NANA
llama.cppV10040961111Local GPU DFlash+30.7%69.7%43.4NANA
llama.cppV10040963434Local GPU DFlash+39.8%78.7%46.4NANA
llama.cppV10040961111Local GPU DFlash+16.1%47.4%38.5NANA
llama.cppV10040963434Local GPU DFlash+38.4%62.0%45.9NANA
llama.cppV10040961111Local GPU DFlash-13.3%33.1%28.8NANA
llama.cppV10040963434Local GPU DFlash+8.2%44.9%35.9NANA
llama.cppV10040961111Local GPU DFlash-13.1%33.1%28.8NANA
llama.cppV10040963434Local GPU DFlash+7.7%44.9%35.7NANA
llama.cppV10040961111Local GPU MTP+47.0%74.7%48.8NANA
llama.cppV10040963434Local GPU MTP+46.2%74.7%48.5NANA
llama.cppV10040961111Local GPU MTP+27.2%68.9%70.141.6NANA
llama.cppV10040961111Local GPU MTP+32.4%74.7%69.643.3NANA
llama.cppV100(2)40963434Local GPU MTP+37.2%78.6%184.244.8NANA
llama.cppV100(2)40963434Local GPU MTP+23.8%74.7%181.940.4NANA
llama.cppV10040961111Local GPU MTP+28.1%74.7%65.642.4NANA
llama.cppV100(2)40963434Local GPU MTP+27.1%74.7%178.342.0NANA
llama.cppV10040961111Local GPU MTP-18.2%73.1%25.626.1NANA
llama.cppV100(2)40963434Local GPU MTP-18.8%73.4%93.325.9NANA
llama.cppV10040961111Local GPU MTP+38.4%87.3%19.646.5NANA
llama.cppV100(2)40963434Local GPU MTP+38.1%89.1%64.446.4NANA
llama.cppV10040961111Local GPU MTP+30.6%62.0%30.543.9NANA
llama.cppV100(2)40963434Local GPU MTP+35.4%64.9%72.045.5NANA
llama.cppV10040961111Local GPU MTP+24.3%69.3%25.541.8NANA
llama.cppV100(2)40963434Local GPU MTP+32.9%73.7%63.544.7NANA
llama.cppV10040961111Local GPU MTP+31.3%57.6%28.644.1NANA
llama.cppV100(2)40963434Local GPU MTP+28.8%57.0%76.543.3NANA
llama.cppV10040961111Local GPU MTP+1.1%43.0%20.034.0NANA
llama.cppV100(2)40963434Local GPU MTP+7.2%46.5%56.936.0NANA
llama.cppV10040961111Local GPU MTP+43.3%74.7%19.749.1NANA
llama.cppV100(2)40963434Local GPU MTP+29.1%74.7%73.344.2NANA
llama.cppV100(2)409603090630906Local GPU MTP+63.6%88.1%552.545.624%NA
llama.cppV100(2)409603090630906Local GPU MTP+65.8%88.1%543.246.424%NA
llama.cppV100327683090630906Local GPU MTP+62.7%88.1%552.945.224%NA
llama.cppV100131072111781111781Local GPU MTP-11.3%429.424.819%NA
llama.cppV100131072112941112941Local GPU MTP+17.6%427.932.919%NA
llama.cppV100131072116943116943Local GPU MTP-0.3%420.027.918%NA
llama.cppV100131072118103118103Local GPU MTP-13.8%418.424.118%NA
llama.cppV100131072111781111781Local GPU MTP+4.2%430.025.919%NA
llama.cppV100131072112941112941Local GPU MTP+33.0%428.533.019%NA
llama.cppV100131072116943116943Local GPU MTP+12.6%419.927.918%NA
llama.cppV100131072118103118103Local GPU MTP-2.5%418.324.218%NA
llama.cppV100131072111781111781Local CPU MTP-2.0%423.629.019%NA
llama.cppV100131072111781111781Local CPU MTP-3.8%430.128.519%NA
llama.cppV100131072111781111781Local CPU MTP+5.7%365.929.616%NA
llama.cppV100131072111781111781Local GPU MTP69.2%417.818%NA
llama.cppV100131072116943116943Local GPU MTP72.0%412.918%NA
llama.cppV100131072112941112941Local GPU MTP90.9%226.310%NA
llama.cppV100131072118103118103Local GPU MTP55.2%223.210%NA

Qwen3.6

Size
28B
Quant
Q4_K_M
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cppV100TODO128305.131.8NA59%
llama.cpp/poolsideV100(2)13107212832768225.927.5NA
llama.cpp/poolsideV100(2)13107212865536190.123.3NA
llama.cpp/poolsideV100(2)131072128131072139.017.6NA

Qwen3.6

Size
28B
Quant
UD-Q4_K_XL
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cpp/lagunaV100409619119.125.9NA52%
llama.cpp/lagunaV100409619Local GPU MTP+52.2%89.6%84.339.4NANA

Muse Glimmer 30B

Size
30B
Quant
K-Quant-17GB
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cppV10040961182.637.8NA70%
llama.cppV1004096512845.238.740%72%
llama.cppV10032768128296.338.7NA72%
llama.cppV1004096512?824.539%
llama.cppV100327681282048294.638.2NA71%
llama.cppV100327681288192287.238.0NA
llama.cppV1003276812832768273.436.8NA
llama.cppV100409611Local GPU DFlash-20.8%15.0%79.630.0NANA
llama.cppV100409611Local GPU DFlash+20.1%78.9%79.345.4NANA
llama.cppV100409611Local GPU DFlash+19.7%64.0%79.745.3NANA
llama.cppV100409611Local GPU DFlash-1.3%41.8%78.937.3NANA
llama.cppV100409611Local GPU DFlash-16.5%25.5%80.031.6NANA

Muse Glimmer 30B

Size
30B
Quant
K-Quant-Dynamic
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cppV10040961179.632.1NA70%
llama.cppV1004096512897.533.043%72%
llama.cppV100409611Local GPU DFlash-17.6%12.8%75.826.4NANA
llama.cppV100409611Local GPU DFlash+28.1%82.6%74.841.1NANA
llama.cppV100409611Local GPU DFlash+34.8%66.1%73.843.3NANA
llama.cppV100409611Local GPU DFlash+16.2%43.5%78.537.3NANA
llama.cppV100409611Local GPU DFlash-7.2%25.4%77.729.8NANA

Gemma 4

Size
31B
Quant
UD-Q4_K_XL
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cppCPUTODO1284.92.6NANA
llama.cppCPU4096TODO2.5NANA
llama.cppCPU4096TODOLocal CPU MTP+35.4%77.5%3.4NANA
llama.cpp/lagunaV100TODO128258.935.6NA68%
llama.cpp/poolsideV100409620144.233.0NA63%
llama.cpp/poolsideCPU + V100409620Local CPU MTP-4.5%63.6%139.131.5NANA

Laguna XS 2.1

Size
33B / 3B
Quant
Q4_K_M
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cpp/poolsideV100409611180.177.3NANA
llama.cpp/poolsideV1004096128422.095.9NANA
llama.cpp/poolsideV100409611Local GPU DFlash+35.3%50.0%275.6104.6NANA

Qwen3.6 35B-A3B

Size
35B / 3B
Quant
Q4_K_M
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cpp/poolsideV100(2)1024512486.580.32%NA
llama.cpp/poolsideV10040961176.474.7NANA
llama.cpp/poolsideV1001024512488.179.22%NA
llama.cpp/poolsideV100409611Local GPU MTP+29.8%84.1%77.096.9NANA
llama.cpp/poolsideV100409611Local GPU MTP+35.9%59.6%76.8101.5NANA
llama.cpp/poolsideV100409611Local GPU MTP+25.6%46.2%74.993.8NANA

Dolphin 2.2

Size
70B
Quant
Q3_K_M
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
PipeInfer/lanCPU + V1001024TODO2.4NANA
PipeInfer/localCPU + V1001024TODO2.4NANA
llama.cpp/poolsideV100(2)32768128140.65.7NA21%
llama.cpp/poolsideV100327681288192127.911.5NA
llama.cpp/poolsideV1003276812832768104.210.4NA
llama.cpp/poolsideV100(2)1024512138.68.416%31%
llama.cpp/poolsideV1001024512106.66.412%24%

Qwen2.5

Size
72B
Quant
Q4_K_M
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
prima.cppCPU + P40TODOTODO1.52.0NANA
prima.cppCPU + P401024TODOLocal CPU 3B draft+6.1%66.0%2.1NANA
prima.cppCPU + P401024TODORemote CPU 3B draft+13.9%66.0%2.2NANA
prima.cppCPU + V1001024190.32.3NANA
prima.cppCPU + V1001024197.54.1NANA
prima.cppCPU + V1001024197.02.8NANA
llama.cpp/poolsideV100(2)32768128127.58.3NA
llama.cpp/poolsideV100(2)327681288192118.08.6NA
llama.cpp/poolsideV100(2)327681283276895.86.6NA
llama.cpp/poolsideV100(2)1024512117.88.314%
llama.cpp/poolsideV100(2)102451227.32.6NANA
prima.cppCPU + V100102419Local CPU 3B draft+11.1%68.6%3.62.6NANA
prima.cppCPU + V100102419Remote CPU 3B draft-3.1%66.0%4.0NANA
prima.cppCPU + V100102419Remote CPU 3B draft+39.9%66.0%6.63.9NANA
prima.cppCPU + V100102419Remote CPU 3B draft+14.6%68.6%4.02.7NANA

Qwen3-Coder-Next 80B-A3B

Size
80B / 3B
Quant
Q4_K_M
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cppV100(2)1024512270.579.61%NA
llama.cppCPU + V100102451252.228.6NANA

gpt-oss-120B

Size
117B / 5B
Quant
MXFP4
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cppV100(2)1024512627.499.95%NA
llama.cppCPU + V100102451227.29.5NANA

Laguna S 2.1

Size
118B / 8B
Quant
Q4_K_M
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cpp/poolsideCPU + V10040961118.512.9NANA
llama.cpp/poolsideCPU + V10040961131.018.5NANA
llama.cpp/poolsideCPU + V10040964096123.717.6NANA
llama.cpp/poolsideCPU + V10040961286.512.0NANA
llama.cpp/poolsideCPU + V100TODO013.2NANA
llama.cpp/poolsideCPU + V1004096119.813.4NANA
llama.cpp/lagunaCPU + V10040961289.67.7NANA
llama.cpp/poolsideCPU + V10040961114.512.5NANA
llama.cpp/poolsideCPU + V100(2)131072128221170.612.6NANA
llama.cpp/poolsideCPU + V100(2)13107229546180.719.3NANA
llama.cpp/poolsideCPU + V100(2)13107262067179.218.6NANA
llama.cpp/poolsideCPU + V100(2)819212862.030.6NANA
llama.cpp/poolsideCPU + V100(2)8192128409665.928.0NANA
llama.cpp/poolsideCPU + V100(2)TODO032.5NANA
llama.cpp/poolsideCPU + V100(2)TODO032.2NANA
llama.cpp/poolsideCPU + V100(2)40961130.225.4NANA
llama.cpp/poolsideCPU + V100(2)4096032.1NANA
llama.cpp/poolsideCPU + V100(2)4096032.2NANA
llama.cpp/poolsideCPU + V100(2)4096018.1NANA
llama.cpp/poolsideCPU + V100(2)1310721283276862.223.6NANA
llama.cpp/poolsideCPU + V100(2)1310721286553655.517.6NANA
llama.cpp/poolsideCPU + V100(2)13107212813107246.815.4NANA
llama.cpp/poolsideV100(3)4096171120.237.7NANA
llama.cpp/poolsideV100(3)40961126.637.3NANA
llama.cpp/poolsideV100(3)TODO038.9NANA
llama.cpp/poolsideV100(3)TODO038.9NANA
llama.cpp/poolsideV100(3)4096133143.339.1NANA
llama.cpp/poolsideV100(3)TODO026.4NANA
llama.cpp/poolsideV100(3)TODO025.6NANA
llama.cpp/poolsideCPU + V100409611Local GPU DFlash+2.0%66.7%20.213.2NANA
llama.cpp/poolsideCPU + V100409611Local GPU DFlash-6.4%66.7%28.817.3NANA
llama.cpp/poolsideCPU + V100409611Local GPU DFlash-41.2%18.4%16.67.4NANA

Mistral Small 4 119B 2603

Size
119B
Quant
UD-IQ3_XXS
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cppV100(2)1024512422.675.1NA
llama.cppV100(2)1024512271.535.2NA
llama.cppV100(2)1024512?NA
llama.cppCPU + V100102451244.212.0NANA
llama.cppCPU + V100TODO12832768110.148.9NANA
llama.cppV100(2)TODO12813107224.524.9NANA

Qwen3.5 122B-A10B

Size
125B / 10B
Quant
Q3_K_M
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cppV100(2)1024512240.144.34%NA
llama.cppCPU + V100102451227.38.7NANA

Qwen3.5 122B-A10B

Size
125B / 10B
Quant
Q5_K_M
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cppCPU + V100(2)TODO015.6NANA
llama.cppCPU + V100(2)TODO011.7NANA
llama.cppV100(3)TODO028.1NANA
llama.cppV100(3)TODO027.6NANA

Step 3.5 Flash

Size
197B / 11B
Quant
Q4_K_M
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cppV100(4)327680?NANA
llama.cppV100(4)327682469924699?245.727.34%NA
llama.cppV100(4)327682469924699?373.726.87%NA
llama.cppV100(4)327682469924699?27.2NA
llama.cppV100(4)131072109941109941?155.123.53%NA
llama.cppV100(4)131072109941109941?262.023.35%NA
llama.cppV100(4)327680Local CPU MTPNANA

MiniMax M2.5

Size
229B / 10B
Quant
Q3_K_M
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cppV100(4)1310720?NANA
llama.cppV100(4)1310720?NANA
llama.cppV100(4)327682382723827?230.520.94%NA
llama.cppV100(4)327682382723827?352.121.36%NA
llama.cppV100(4)131072104561104561?100.26.92%NA
llama.cppV100(4)131072104561104561?183.87.23%NA

DeepSeek V4 Flash 0731

Size
284B / 13B
Quant
UD-Q8_K_XL
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
llama.cppCPU + V100(2)409651215.94.1NANA
llama.cppCPU + V100(2)409651215.93.6NANA
llama.cppCPU + V100(2)409651210.64.1NANA

GLM-5.2

Size
744B / 40B
Quant
INT4 g64 + INT8 MTP
CfgBackendHardwareWindowPromptTest depthSpeculativePrefill tok/sDecode tok/sPrefill peakDecode peak
EnabledSpeedupAccept
ColibriCPU + V100(2)4096170.6NANA
ColibriCPU + V100(2)4096170.8NANA
ColibriCPU + V100(2)4096170.1NANA
ColibriCPU + V100(2)4096170.1NANA
ColibriCPU + V1004096170.6NANA
ColibriCPU + V1004096170.6NANA
ColibriCPU + V1004096170.4NANA
ColibriCPU + V1004096170.4NANA
ColibriCPU + V1004096170.3NANA
ColibriCPU + V100(2)4096170.2NANA
ColibriCPU + V100(2)4096170.2NANA
ColibriCPU + V100(2)4096170.2NANA
ColibriCPU + V1004096170.2NANA
ColibriCPU + V100(2)4096170.7NANA
ColibriCPU + V100(2)4096170.4NANA
ColibriCPU + V100(2)409617?NANA
ColibriCPU + V100(2)409617?NANA
ColibriCPU + V100(2)4096170.7NANA
ColibriCPU + V100(2)409617?NANA
ColibriCPU + V100(2)4096171.2NANA
ColibriCPU + V100(2)409617Local GPU MTP-33.5%0.4NANA
ColibriCPU + V100409617Local GPU MTP-77.8%0.1NANA

Transport and other workloads

Select any result row for complete configuration details. Non-model workloads use task-specific metrics. Collective rows show median out-of-place values; tensor proof rows show one eight-block step; host-device transfer rows show the median rate of an 8 GiB copy in each direction.

CfgWorkloadBackendHostsAcceleratorsInterconnectValiditySamples4 KiB us16 MiB GB/sStep msSteps/s8 GiB H2D GB/s8 GiB D2H GB/sSource
Eight-block tensor-parallel MLPCUDA/NCCLr720Tesla PG500-216 (V100 32GB)Local V100 HBMvalid52.471404.771distributed/results/20260811T055619Z-v100-p40-tp-control
Eight-block tensor-parallel MLPCUDA/NCCLr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1 over direct ConnectX-3 40 GbE; host stagingvalid53.344299.072distributed/results/20260811T055619Z-v100-p40-tp-roce-host
Eight-block tensor-parallel MLPCUDA/NCCLr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1 over direct ConnectX-3 40 GbE; GPUDirect RDMAvalid53.342299.244distributed/results/20260811T055619Z-v100-p40-tp-roce-gdr
GLM-5.2Colibrir720valid5colibri/results/20260814T212000Z-final-sse41-grouped-int4-t40/README.md
GLM-5.2Colibrir720valid5colibri/results/20260814T212000Z-final-sse41-grouped-int4-t40/README.md
GLM-5.2Colibrir720valid5colibri/results/20260814T212100Z-final-sse41-grouped-int4-t20/README.md
GLM-5.2Colibrir720valid5colibri/results/20260814T212100Z-final-sse41-grouped-int4-t20/README.md
GLM-5.2Colibrir720Tesla PG500-216 (V100 32GB)valid5colibri/results/20260815T002617Z-g64-warp-v100/README.md
GLM-5.2Colibrir720Tesla PG500-216 (V100 32GB)valid5colibri/results/20260815T002617Z-g64-warp-v100/README.md
GLM-5.2Colibrir720valid5colibri/results/20260830T212230Z-ivybridge-grouped-int4/README.md
GLM-5.2Colibrir720valid5colibri/results/20260830T212230Z-ivybridge-grouped-int4/README.md
GLM-5.2Colibrir720valid5colibri/results/20260830T212251Z-ivybridge-grouped-int4/README.md
GLM-5.2Colibrir720valid5colibri/results/20260830T212251Z-ivybridge-grouped-int4/README.md
GPU peer transferdevice-peer-xferr720Tesla PG500-216 (V100 32GB)PCIe 3.0 x16 pair, 0000:05:00.0 to 0000:42:00.0; SYS path without peer accessvalid517.9010.65pcie/results/20260821T233825Z-r720-nccl-p2p
GPU peer transferdevice-peer-xferr720Tesla PG500-216 (V100 32GB)PCIe 3.0 x16 pair, 0000:05:00.0 to 0000:42:00.0; SYS path without peer accessvalid520.576.04pcie/results/20260821T233825Z-r720-nccl-p2p
GPU peer transferdevice-peer-xferr720Tesla PG500-216 (V100 32GB)PCIe 3.0 x16 pair, 0000:05:00.0 to 0000:42:00.0; SYS path without peer accessvalid517.3510.59pcie/results/20260821T233825Z-r720-nccl-p2p
GPU peer transferdevice-peer-xferr720Tesla PG500-216 (V100 32GB)PCIe 3.0 x16 pair, 0000:05:00.0 to 0000:42:00.0; SYS path without peer accessvalid518.916.22pcie/results/20260821T233825Z-r720-nccl-p2p
GPU peer transferdevice-peer-xferr720Tesla PG500-216 (V100 32GB)PCIe 3.0 x16 pair, 0000:42:00.0 to 0000:05:00.0; SYS path without peer accessvalid515.5210.51pcie/results/20260821T233825Z-r720-nccl-p2p
GPU peer transferdevice-peer-xferr720Tesla PG500-216 (V100 32GB)PCIe 3.0 x16 pair, 0000:42:00.0 to 0000:05:00.0; SYS path without peer accessvalid520.506.04pcie/results/20260821T233825Z-r720-nccl-p2p
GPU peer transferdevice-peer-xferr720Tesla PG500-216 (V100 32GB)PCIe 3.0 x16 pair, 0000:42:00.0 to 0000:05:00.0; SYS path without peer accessvalid516.359.83pcie/results/20260821T233825Z-r720-nccl-p2p
GPU peer transferdevice-peer-xferr720Tesla PG500-216 (V100 32GB)PCIe 3.0 x16 pair, 0000:42:00.0 to 0000:05:00.0; SYS path without peer accessvalid518.756.10pcie/results/20260821T233825Z-r720-nccl-p2p
NCCL Tests FP16 all-reduceNCCL Testsr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40TCP over direct ConnectX-3 40 GbEvalid578.141.71distributed/results/20260809T081110Z-v100-p40-tcp-40gbe
NCCL Tests FP16 all-reduceNCCL Testsr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1 over direct ConnectX-3 40 GbE; host stagingvalid523.013.83distributed/results/20260809T081110Z-v100-p40-roce-host
NCCL Tests FP16 all-reduceNCCL Testsr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40PRRTE control launch toward direct RoCE linkinvalid6distributed/results/20260810T211644Z-v100-p40-roce-gdr-launch-screens
NCCL Tests FP16 all-reduceNCCL Testsr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1 over direct ConnectX-3 40 GbE; GPUDirect RDMAvalid127.25distributed/results/20260810T211644Z-v100-p40-roce-gdr-launch-screens/nccl-gdr-screen-20260810T211644Z-7
NCCL Tests FP16 all-reduceNCCL Testsr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1 over direct ConnectX-3 40 GbE; host stagingvalid522.243.87distributed/results/20260810T211644Z-v100-p40-roce-host-patched
NCCL Tests FP16 all-reduceNCCL Testsr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1 over direct ConnectX-3 40 GbE; host stagingvalid522.823.84distributed/results/20260810T225836Z-v100-p40-roce-host-control
NCCL Tests FP16 all-reduceNCCL Testsr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1 over direct ConnectX-3 40 GbE; PHB GPUDirect RDMAvalid521.752.31distributed/results/20260810T211644Z-v100-p40-roce-gdr-phb
NCCL Tests FP16 all-reduceNCCL Testsr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1; direct GPU receives and host-staged sendsvalid522.254.31distributed/results/20260810T225836Z-v100-p40-roce-gdr-write-only
NCCL Tests FP16 all-reduceNCCL Testsr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1; direct GPU receives and sendsvalid521.812.31distributed/results/20260810T225836Z-v100-p40-roce-gdr-read-write
NCCL Tests FP16 all-reduceNCCL Testsr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1 port 2; host stagingvalid523.143.84distributed/results/20260814T154150Z-host-rail0
NCCL Tests FP16 all-reduceNCCL Testsr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1 port 1; host stagingvalid522.503.84distributed/results/20260814T154150Z-host-rail1
NCCL Tests FP16 all-reduceNCCL Testsr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1 fused ports; host stagingvalid528.843.84distributed/results/20260814T154150Z-host-dual
NCCL Tests FP16 all-reduceNCCL Testsr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1 port 2; direct receive and staged sendvalid521.844.31distributed/results/20260814T154150Z-hybrid-rail0
NCCL Tests FP16 all-reduceNCCL Testsr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1 port 1; direct receive and staged sendvalid522.124.26distributed/results/20260814T154150Z-hybrid-rail1
NCCL Tests FP16 all-reduceNCCL Testsr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1 fused ports; direct receive and staged sendvalid527.994.90distributed/results/20260814T154150Z-hybrid-dual
NCCL Tests FP16 all-reduceNCCL Testsr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1 port 2; full GDRvalid532.262.31distributed/results/20260814T154150Z-full-rail0
NCCL Tests FP16 all-reduceNCCL Testsr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1 port 1; full GDRvalid521.602.31distributed/results/20260814T154150Z-full-rail1
NCCL Tests FP16 all-reduceNCCL Testsr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1 fused ports; full GDRvalid529.202.30distributed/results/20260814T154150Z-full-dual
NCCL Tests FP16 all-reduceNCCL Testsr720Tesla PG500-216 (V100 32GB)PCIe 3.0 x16 pair; NCCL selected SHM/direct/directvalid520.627.20pcie/results/20260821T233825Z-r720-nccl-p2p
NCCL Tests FP16 all-reduceNCCL Testsr720Tesla PG500-216 (V100 32GB)PCIe 3.0 x16 pair; NCCL selected SHM/direct/directvalid520.707.01pcie/results/20260821T233825Z-r720-nccl-p2p
Network transport bandwidth preflightiperf3r720,r720xdTCP over ConnectX-3 port 2valid5distributed/results/20260814T153301Z-dual-rail-preflight
Network transport bandwidth preflightiperf3r720,r720xdTCP over ConnectX-3 port 1valid5distributed/results/20260814T153301Z-dual-rail-preflight
Network transport bandwidth preflightiperf3r720,r720xdConcurrent TCP over both ConnectX-3 portsvalid5distributed/results/20260814T153301Z-dual-rail-preflight
PCIe host-device transferhost-device-xferr720Tesla PG500-216 (V100 32GB)PCIe 3.0 x16 at 0000:05:00.0; NUMA node 0valid52.132.06pcie/results/20260821T104000Z-r720-pcie-x16
PCIe host-device transferhost-device-xferr720Tesla PG500-216 (V100 32GB)PCIe 3.0 x16 at 0000:05:00.0; NUMA node 0valid511.4312.48pcie/results/20260821T104000Z-r720-pcie-x16
PCIe host-device transferhost-device-xferr720Tesla PG500-216 (V100 32GB)PCIe 3.0 x16 at 0000:05:00.0; NUMA node 0valid53.313.24pcie/results/20260821T104000Z-r720-pcie-x16
PCIe host-device transferhost-device-xferr720Tesla PG500-216 (V100 32GB)PCIe 3.0 x16 at 0000:05:00.0; NUMA node 0valid511.7912.95pcie/results/20260821T104000Z-r720-pcie-x16
PCIe host-device transferhost-device-xferr720Tesla PG500-216 (V100 32GB)PCIe 3.0 x16 at 0000:42:00.0; NUMA node 1valid52.011.94pcie/results/20260821T104000Z-r720-pcie-x16
PCIe host-device transferhost-device-xferr720Tesla PG500-216 (V100 32GB)PCIe 3.0 x16 at 0000:42:00.0; NUMA node 1valid511.4112.56pcie/results/20260821T104000Z-r720-pcie-x16
PCIe host-device transferhost-device-xferr720Tesla PG500-216 (V100 32GB)PCIe 3.0 x16 at 0000:42:00.0; NUMA node 1valid53.443.26pcie/results/20260821T104000Z-r720-pcie-x16
PCIe host-device transferhost-device-xferr720Tesla PG500-216 (V100 32GB)PCIe 3.0 x16 at 0000:42:00.0; NUMA node 1valid512.1013.01pcie/results/20260821T104000Z-r720-pcie-x16
perftest CUDA RDMA bandwidthperftestr720,r720xdRoCE v1 RDMA WRITE; host to hostvalid5distributed/results/20260810T225836Z-cuda-rdma-write-host-host
perftest CUDA RDMA bandwidthperftestr720,r720xdTesla PG500-216 (V100 32GB)RoCE v1 RDMA WRITE; host to V100 CUDA memoryvalid5distributed/results/20260810T225836Z-cuda-rdma-write-host-to-v100
perftest CUDA RDMA bandwidthperftestr720,r720xdTesla P40RoCE v1 RDMA WRITE; P40 CUDA memory to hostvalid5distributed/results/20260810T225836Z-cuda-rdma-write-p40-to-host
perftest CUDA RDMA bandwidthperftestr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1 RDMA WRITE; P40 CUDA memory to V100 CUDA memoryvalid5distributed/results/20260810T225836Z-cuda-rdma-write-p40-to-v100
perftest CUDA RDMA bandwidthperftestr720,r720xdRoCE v1 RDMA READ; host to hostvalid5distributed/results/20260810T225836Z-cuda-rdma-read-host-host
perftest CUDA RDMA bandwidthperftestr720,r720xdTesla PG500-216 (V100 32GB)RoCE v1 RDMA READ; V100 CUDA memory to hostvalid5distributed/results/20260810T225836Z-cuda-rdma-read-v100-to-host
perftest CUDA RDMA bandwidthperftestr720,r720xdTesla P40RoCE v1 RDMA READ; host to P40 CUDA memoryvalid5distributed/results/20260810T225836Z-cuda-rdma-read-host-to-p40
perftest CUDA RDMA bandwidthperftestr720,r720xdTesla PG500-216 (V100 32GB),Tesla P40RoCE v1 RDMA READ; V100 CUDA memory to P40 CUDA memoryvalid5distributed/results/20260810T225836Z-cuda-rdma-read-v100-to-p40
perftest CUDA RDMA bandwidthperftestr720,r720xdRoCE v1 RDMA WRITE over ConnectX-3 port 2valid5distributed/results/20260814T153301Z-verbs-host-rail0
perftest CUDA RDMA bandwidthperftestr720,r720xdRoCE v1 RDMA WRITE over ConnectX-3 port 1valid5distributed/results/20260814T153301Z-verbs-host-rail1
perftest CUDA RDMA bandwidthperftestr720,r720xdConcurrent RoCE v1 RDMA WRITEs over both portsvalid5distributed/results/20260814T154150Z-dual-rail-roce.md