(无标题)
very interesting table from deepseek v3.2 that compares the output token count on different benchmarks, dsv3.2 speciale version thinks much more than any other model, BUT since they are using sparse attention the inference cost will still be ok? https://t.co/qc0iu4nx5A