Dataset Viewer
Auto-converted to Parquet Duplicate
data_source
string
prompt
list
ability
string
reward_model
dict
extra_info
dict
code_taco
[ { "role": "system", "content": "You are an expert Python programmer. You will be given a question (problem specification) and will generate a correct Python program that matches the specification and passes all tests.\n\n### Format: Read the inputs from stdin solve the problem and write the answer to stdout...
code
{ "style": "rule", "extraction_method": null, "ground_truth": "{\"inputs\": [\"abc\\n3\\n1 2 2 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1\\n\", \"mmzhr\\n3\\n443 497 867 471 195 670 453 413 579 466 553 881 847 642 269 996 666 702 487 209 257 741 974 133 519 453\\n\", \"ajeeseerqnpaujubmajpibxrccazaawetywxmifze...
{ "id": "TACO_9e61bd10-5471-4b7f-af5e-374400d2d007", "lower_pass_rate": 0, "upper_pass_rate": 0.5 }
code_primeintellect
[ { "role": "system", "content": "You are an expert Python programmer. You will be given a question (problem specification) and will generate a correct Python program that matches the specification and passes all tests.\n\n### Format: Read the inputs from stdin solve the problem and write the answer to stdout...
code
{ "style": "rule", "extraction_method": null, "ground_truth": "{\"inputs\": [\"RYBGRYBGR\\n\", \"!RGYB\\n\", \"!!!!YGRB\\n\", \"!GB!RG!Y!\\n\", \"RYBG\\n\", \"!Y!!!Y!!G!!!G!!B!!R!!!!B!!!!!Y!!G!R!!!!!!!!!!!!B!!!!GY!B!!!!!YR!G!!!!!!B!Y!B!!!!!!R!G!!!!!!!G!R!!!!B\\n\", \"!R!GBRYG!RYGB!!G!!YG!!Y!!\\n\", \"RBGYRBGYRBGY...
{ "id": "PRIMEINTELLECT_27caff35-0e93-4779-96c7-6112eebe11f8", "lower_pass_rate": 0, "upper_pass_rate": 0.25 }
code_taco
[ { "role": "system", "content": "You are an expert Python programmer. You will be given a question (problem specification) and will generate a correct Python program that matches the specification and passes all tests.\n\n### Format: Read the inputs from stdin solve the problem and write the answer to stdout...
code
{ "style": "rule", "extraction_method": null, "ground_truth": "{\"inputs\": [\"6\\n6 4 2 7 2 7\\n3\\n2 3 6\\n1 3 4\\n1 1 6\\n\", \"4\\n5 5 2 3\\n10\\n1 2 4\\n2 1 4\\n1 1 1\\n2 1 4\\n2 1 2\\n1 1 1\\n1 3 3\\n1 1 3\\n1 4 4\\n1 2 2\\n\", \"4\\n2 2 3 6\\n9\\n2 2 3\\n1 1 3\\n2 2 3\\n2 2 3\\n2 2 2\\n1 1 3\\n1 1 3\\n2 1 ...
{ "id": "TACO_b6f07cfc-cabe-4b80-9589-357a95953dc0", "lower_pass_rate": 0.625, "upper_pass_rate": 1 }
code_primeintellect
[{"role":"system","content":"You are an expert Python programmer. You will be given a question (prob(...TRUNCATED)
code
{"style":"rule","extraction_method":null,"ground_truth":"{\"inputs\": [\"3 1\\n\", \"4 3\\n\", \"2 1(...TRUNCATED)
{"id":"PRIMEINTELLECT_739a699e-a25e-4ce0-92e0-85e91ef01a28","lower_pass_rate":0.0,"upper_pass_rate":(...TRUNCATED)
code_taco
[{"role":"system","content":"You are an expert Python programmer. You will be given a question (prob(...TRUNCATED)
code
{"style":"rule","extraction_method":null,"ground_truth":"{\"inputs\": [\"7\\n\\n1 1\\n3 3\\n2 2\\n\\(...TRUNCATED)
{ "id": "TACO_7acf251c-05f4-4148-a24e-c702e9115cec", "lower_pass_rate": 0, "upper_pass_rate": 0.625 }
code_primeintellect
[{"role":"system","content":"You are an expert Python programmer. You will be given a question (prob(...TRUNCATED)
code
{"style":"rule","extraction_method":null,"ground_truth":"{\"inputs\": [\"4\\n0101\\n1000\\n1111\\n01(...TRUNCATED)
{"id":"PRIMEINTELLECT_d8841f66-b019-4398-b93f-34a869a19f8c","lower_pass_rate":0.0,"upper_pass_rate":(...TRUNCATED)
code_primeintellect
[{"role":"system","content":"You are an expert Python programmer. You will be given a question (prob(...TRUNCATED)
code
{"style":"rule","extraction_method":null,"ground_truth":"{\"inputs\": [\"2\", \"3\", \"4\", \"5\", \(...TRUNCATED)
{"id":"PRIMEINTELLECT_cfc3ef0f-e251-4474-90b1-a627117d8e9f","lower_pass_rate":0.125,"upper_pass_rate(...TRUNCATED)
code_primeintellect
[{"role":"system","content":"You are an expert Python programmer. You will be given a question (prob(...TRUNCATED)
code
{"style":"rule","extraction_method":null,"ground_truth":"{\"inputs\": [\"2\\n1 3 2 4\\n\", \"1\\n3 3(...TRUNCATED)
{"id":"PRIMEINTELLECT_5e6a7525-1ba9-43cc-b717-469c305cc304","lower_pass_rate":0.0,"upper_pass_rate":(...TRUNCATED)
code_primeintellect
[{"role":"system","content":"You are an expert Python programmer. You will be given a question (prob(...TRUNCATED)
code
{"style":"rule","extraction_method":null,"ground_truth":"{\"inputs\": [\"4 3\\n-1 0 3\\n0 0 3\\n1 0 (...TRUNCATED)
{"id":"PRIMEINTELLECT_33e3a605-3602-4fb8-98c5-22005a1cca3c","lower_pass_rate":0.0,"upper_pass_rate":(...TRUNCATED)
code_taco
[{"role":"system","content":"You are an expert Python programmer. You will be given a question (prob(...TRUNCATED)
code
{"style":"rule","extraction_method":null,"ground_truth":"{\"inputs\": [\"4\\n1 1\\n999999999 1000000(...TRUNCATED)
{ "id": "TACO_64807842-b216-4dfd-a5fc-eb76086d11ff", "lower_pass_rate": 0, "upper_pass_rate": 0.125 }
End of preview. Expand in Data Studio

FinalMix2

A multi-task code reinforcement-learning dataset mixture in the verl RL prompt format. It pairs a code-generation split with a suite of auxiliary code-understanding tasks so the same corpus can drive three training regimes from one repo. It is the V3-dedupe successor to OctoReasoner/FinalMix (see Relationship to FinalMix (v1)).

Splits

Split Rows Contents Use
train_no_aux 9,693 code-generation only RL on code gen alone
train_aux_cascade 25,538 all 15,845 auxiliary rows first, then the 9,693 code rows appended (order preserved) cascade / curriculum RL (aux β†’ code)
train_aux_multitask 25,538 the same code + aux rows concatenated and shuffled (seed=42) mixed multi-task RL
validation 481 held-out code-generation problems eval
test 175 LiveCodeBench-v6 problems eval

The three training splits are built from the same underlying rows β€” they differ only in which tasks are included and in what order β€” so they form a controlled three-way comparison:

  1. train_no_aux β€” code generation only.
  2. train_aux_cascade β€” auxiliary tasks then code, for cascade RL.
  3. train_aux_multitask β€” code and auxiliary tasks interleaved, for mixed multi-task RL.
from datasets import load_dataset

code_only  = load_dataset("OctoReasoner/FinalMix2", split="train_no_aux")
cascade    = load_dataset("OctoReasoner/FinalMix2", split="train_aux_cascade")
multitask  = load_dataset("OctoReasoner/FinalMix2", split="train_aux_multitask")
val        = load_dataset("OctoReasoner/FinalMix2", split="validation")
test       = load_dataset("OctoReasoner/FinalMix2", split="test")

Code split (9,693)

A more liberal ("V3") deduplication of the source code pools, rebalanced away from the contest-heavy v1 mix toward PrimeIntellect:

Source Rows Share
code_primeintellect 5,241 54.1%
code_contests_o 2,538 26.2%
code_taco 1,721 17.8%
code_lcbv5 193 2.0%

Auxiliary tasks (15,845)

Twelve data_sources spanning ~24 ability sub-tasks that probe code understanding beyond generation:

  • Input/output reasoning β€” code_io_taco, code_functional_identity (predict outputs from inputs / inputs from outputs, direct and MCQ).
  • Complexity β€” code_time_complexity, code_space_complexity, code_cpu_ranking, code_memory_ranking (predict/rank time, space, CPU, memory).
  • Security β€” code_sast_cwe (predict/localize CWE weaknesses).
  • Retrieval β€” code_crp_retrieval (coderpile_retrieval).
  • Localization β€” code_change_localization, code_var_tracing (locate edits; trace variable values).
  • Compilation β€” code_compile_status (predict whether code compiles).
  • Instruction following β€” codeif (verifiable instruction-following, generate & edit).

Schema

Standard verl RL fields:

Field Type Notes
data_source string routes the reward function
prompt list of {role, content} chat-formatted problem
ability string task category
reward_model struct {style, extraction_method, ground_truth, key} scoring spec
extra_info struct {id, lower_pass_rate, upper_pass_rate} per-example metadata

Code-generation rows are scored by executing model output against tests in a sandbox; auxiliary rows are scored by rule / answer extraction against ground_truth.

Relationship to FinalMix (v1)

FinalMix2 rebuilds the code split of OctoReasoner/FinalMix on a more liberal dedupe (9,693 code rows vs. 6,000) and rebalances the source distribution β€” v1 was code_contests_o-dominated (50%), v2 leads with code_primeintellect (54%). The combined training splits grow accordingly (25,538 vs. 22,000). The schema, the auxiliary-task set, and the validation/test eval splits are carried over unchanged from v1.

Downloads last month
260