<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://tytsui.com/feed.xml" rel="self" type="application/atom+xml" /><link href="https://tytsui.com/" rel="alternate" type="text/html" /><updated>2026-10-07T05:01:06+00:00</updated><id>https://tytsui.com/feed.xml</id><title type="html">TSUITUENYUE</title><subtitle>Robotics researcher at UPenn focusing on intention-aware imitation learning, world models, and physics-aware vision/graphics.</subtitle><author><name>Tuen-Yue Tsui</name><email>tytsui@seas.upenn.edu</email><uri>https://tytsui.com/</uri></author><entry><title type="html">Correct, minimal, and all</title><link href="https://tytsui.com/blog/correct-minimal-and-all/" rel="alternate" type="text/html" title="Correct, minimal, and all" /><published>2026-10-07T00:00:00+00:00</published><updated>2026-10-07T00:00:00+00:00</updated><id>https://tytsui.com/blog/correct-minimal-and-all</id><content type="html" xml:base="https://tytsui.com/blog/correct-minimal-and-all/"><![CDATA[<p><strong>TL;DR</strong> Many problems are partially ordered, yet we often reduce minimality and alternatives to preferences. Reinforcement learning with verifiable rewards (RLVR) scores each rollout on its own, so it cannot tell a new minimal answer from a redundant superset of one it already has. This post is about what changes when the RLVR target is actually partially ordered.</p>

<p>During my summer visit to Shanghai, I was given a task: use RL to find which dimensions of a chemical reaction’s condition space are worth searching to accelerate Bayesian optimization (BO) [1, 2].</p>

<p>Take a reaction every high-school chemistry student has met: the Haber process, N<sub>2</sub> + <span class="lining">3</span>H<sub>2</sub> ⇌ <span class="lining">2</span>NH<sub>3</sub>. To make ammonia you choose a catalyst, possibly a few promoters for it, a temperature, a pressure and a ratio of the two gases. The textbook recipe is an iron catalyst at about 450 °C and 200 atm. Take each of these choices as one dimension of the reaction space. BO is very good at searching such a space for the best yield, but every dimension you give it makes the search more expensive. So the question becomes: starting from a baseline recipe, which dimensions do you actually need to turn to reach a good yield?</p>

<p>The task looked trivial: the space is small and its transition model is fully known, so you can use <code class="language-plaintext highlighter-rouge">pymdptoolbox</code>, solve it directly, and succeed 100% of the time.</p>

<p>Seems easy. But is this what we want?</p>

<p>The result it returns simply says: “Include everything.” Of course, including everything IS a correct reaction space: BO will eventually find the right reaction no matter how large the space is. It is also useless, because giving BO a smaller (or better, minimal) space was the whole point.</p>

<h2 id="minimality-is-not-a-preference">Minimality is not a preference</h2>

<p>Why not just add a size penalty? Minimality, meaning that nothing can be removed without losing the outcome, is not a <strong>preference</strong> for small size; it is the very <strong>definition</strong> of a valid answer. When Fritz Haber first got ammonia synthesis to work in 1909, his catalyst was Os, a rare and expensive metal, and the process built on it looked hard to scale. Later, at BASF, Alwin Mittasch tested thousands of catalyst compositions and found that Fe works too, as long as it comes with a couple of promoters such as K<sub>2</sub>O and Al<sub>2</sub>O<sub>3</sub>. This turned Haber’s tabletop demonstration into the industrial Haber–Bosch process. Today, the Haber–Bosch process produces nearly all of the world’s nitrogen used in fertilizers. By preferring a minimum size, we would miss the recipe that feeds the world. In this case, both Os and Fe + K<sub>2</sub>O + Al<sub>2</sub>O<sub>3</sub> are MINIMAL, because each is an irreducible set of conditions that converts nitrogen and hydrogen into ammonia. If you simply ask “what are the irreducible conditions that make the Haber process work?”, the answer should include both.</p>

<h2 id="and-diversity">And diversity?</h2>

<p>What about adding a diversity term or trying soft RL to spread the probability, if we also want alternatives? Alternatives are not a preference here either. If a reaction has one minimal recipe, the best answer is to only recover that recipe. If it has three, the best answer is exactly those three, no more and no less. A diversity reward is blind to this structure.</p>

<p>Moreover, a diversity method needs a metric, such as entropy, to judge whether answers are diverse. Nature imposes no such metric, so using one injects a strong prior that the problem itself does not contain.</p>

<p>In plain words, what we want is a family of recipes in which each recipe is <strong>correct</strong> and <strong>minimal</strong>, and we want to find <strong>all</strong> of them. Finding the minimal answer AND every alternative is in fact a well-studied problem. If you know Boolean functions, you will recognize this as prime-implicant enumeration [3, 4]. Algorithms for prime-implicant enumeration, however, can query a white-box formula as often as they like. Yet our chemists cannot: every query is an hours-long wet-lab experiment, the verifier is a black box, and each new reaction starts from scratch. Amortizing and scaling prime-implicant enumeration are also hard, because the search space has \(2^N\) elements for \(N\) conditions. The question then becomes: given only one bit from a black-box verifier, can RL recover this family, in a way that scales and amortizes? This leads to Minimal Witness Reinforcement Learning (MWRL) [5].</p>

<h2 id="why-this-problem-matters">Why this problem matters</h2>

<p>Because we might be ignoring a bunch of such problems, not just in AI4S but also in RLVR, where the target is often partially ordered. Tired of hours-long reasoning that just hands back an idea you already gave in the prompt, an AI-generated codebase full of unnecessary test files and <span class="lining">SHA-256</span> checks, or a 1,000-page generated proof that keeps circling the same paths under countless assumptions? And not just minimality. Take red teaming, one of the biggest concerns in AI today: a red-teamer rewarded for making a model misbehave can stop at the first jailbreak, but the model is only safe when every jailbreak is patched; a second jailbreak that doesn’t share anything with the first one will survive the patch. If we only ask for one answer, we are left to find the rest by hand. These problems (if formalized correctly) can be addressed by asking “what are the minimal sufficient ways of doing this?” instead of just “\(\text{success reward} - \alpha\,\text{size} + \beta\,\text{diversity} + \gamma\,\text{preference}\)” (which is still incorrect for these problems anyway).</p>

<h2 id="partial-orders-lattices-and-monotonicity">Partial orders, Lattices and Monotonicity</h2>

<p>If you are familiar with these concepts, you may skip this part.</p>

<p>An order is a relation \(\le\) on a set that is reflexive (\(3 \le 3\)), antisymmetric (\(x \le 3\) and \(3 \le x\) only when \(x = 3\)) and transitive (\(3 \le 5\) and \(5 \le 7\) give \(3 \le 7\)) [6]. The order that we are most familiar with is a total order, in which of any two elements, one is below the other. A partial order drops that requirement and lets two elements be incomparable. On the left below, \(a \le b \le c \le d\) is a total order; on the right, \(a \le b \le d\) and \(a \le c \le d\) hold, but neither \(b \le c\) nor \(c \le b\), so \(b\) and \(c\) are incomparable.</p>

<figure class="fig fig--orders">
  <svg viewBox="0 0 300 150" width="300" height="150" role="img" aria-label="A total order a, b, c, d in one chain, and a partial order in which a is below b and c, and b and c are both below d" style="display:block;margin:0 auto;max-width:100%;height:auto">
    <g style="stroke:var(--secondary-text);stroke-width:1.2">
      <line x1="80" y1="128" x2="80" y2="92" /><line x1="80" y1="92" x2="80" y2="56" /><line x1="80" y1="56" x2="80" y2="20" />
      <line x1="220" y1="128" x2="185" y2="74" /><line x1="220" y1="128" x2="255" y2="74" /><line x1="185" y1="74" x2="220" y2="20" /><line x1="255" y1="74" x2="220" y2="20" />
    </g>
    <g style="fill:var(--marble-bg);stroke:var(--accent-color);stroke-width:1.2">
      <circle cx="80" cy="128" r="10" /><circle cx="80" cy="92" r="10" /><circle cx="80" cy="56" r="10" /><circle cx="80" cy="20" r="10" />
      <circle cx="220" cy="128" r="10" /><circle cx="185" cy="74" r="10" /><circle cx="255" cy="74" r="10" /><circle cx="220" cy="20" r="10" />
    </g>
    <g style="fill:var(--primary-text);font:italic 13px 'EB Garamond',Georgia,serif" text-anchor="middle">
      <text x="80" y="132">a</text><text x="80" y="96">b</text><text x="80" y="60">c</text><text x="80" y="24">d</text>
      <text x="220" y="132">a</text><text x="185" y="78">b</text><text x="255" y="78">c</text><text x="220" y="24">d</text>
    </g>
  </svg>
  <figcaption class="fig__caption">Left: a total order. Right: a partial order, in which <i>b</i> and <i>c</i> are incomparable. A line runs up from an element to each element directly above it.</figcaption>
</figure>

<p>In our Haber example, recipes are ordered by “contains”: Os alone lies below Os with K<sub>2</sub>O, since the second recipe contains the first, while Os and Fe + K<sub>2</sub>O + Al<sub>2</sub>O<sub>3</sub> are incomparable because neither contains the other.</p>

<p>All subsets of a ground set \(E\), ordered by inclusion, form the subset lattice \(2^E\). We draw it as a Hasse diagram, with one node per set and a line from \(S\) up to \(T\) when \(T\) adds one element to \(S\) (Figure 1).</p>

<figure class="fig" data-fig="lattice">
  <div class="fig__canvas"></div>
  <figcaption class="fig__caption"><b>Figure 1</b>The 16 recipes over four ingredients, in a toy version of ammonia synthesis where a recipe is correct when it contains Os, or Fe together with both promoters.</figcaption>
</figure>

<p>In this toy, we assume adding an ingredient to a correct recipe never makes it incorrect. We call such a verifier <strong>monotone</strong>. Under monotonicity a correct recipe \(S\) certifies every recipe above it, its up-set \(\uparrow S = \{T \subseteq E : T \supseteq S\}\). The correct recipes are therefore the union of the up-sets of the minimal ones, and the minimal recipes are the lower boundary of that region. Together they form an antichain. The answer we want is this whole boundary, which no single best point can represent. Suppose a building has two doors and we ask how a burglar can get in. “Through the front door” is correct, and minimal, yet a guard who locks only that door leaves the building open: knowing what suffices takes one minimal answer, while knowing what must be blocked takes all of them.</p>

<h2 id="standard-rl-just-cant-solve-it">Standard RL just can’t solve it</h2>

<p>Existing RL methods, from PPO [7], GRPO [8] and RLOO [9, 10] to pass@k [11], soft RL [12] and RL-style samplers such as GFlowNets [13], assign credit locally: each rollout ends with a proposal, a set \(S_i\), the verifier returns one bit \(y_i\), and its reward depends only on \(S_i\) and \(y_i\): other rollouts enter at most through their own bits, never through how their sets relate to \(S_i\).</p>

<p>However, you can’t tell minimality, or whether a proposal is a new alternative, without knowing how it relates to the other proposals in the group. Os, Os + K<sub>2</sub>O and Fe + K<sub>2</sub>O + Al<sub>2</sub>O<sub>3</sub> all count as a success. You can’t tell the difference between these three, and the RL optimizer will eventually end up at one of them. Thus no matter how you change the optimizer, as long as your credit assignment is local, you can’t address this problem [5]. Figure 2 shows where each kind of objective ends up putting its probability.</p>

<figure class="fig" data-fig="mass">
  <div class="fig__canvas"></div>
  <figcaption class="fig__caption"><b>Figure 2</b>Optimal probability of each recipe under four objectives on the lattice of Figure 1. Ringed: the minimal recipes.</figcaption>
</figure>

<h2 id="relational-credit-assignment">Relational credit assignment</h2>

<p>So what should the credit assignment look like? Let’s start from what the problem actually asks for:</p>

\[M = \min_{\subseteq}\,\{S \subseteq E : s(S) = 1\}.\]

<p>Notice that \(\min_{\subseteq}\) compares sets with each other, so the credit has to look at a group of proposals together. Inside a group whose correct proposals are \(W\), the same objective becomes \(\min_{\subseteq} W\): the correct proposals that contain no other correct proposal of the group. We turn this set-valued objective into a credit for each proposal in three steps:</p>

<p><strong>From minimal elements to a region.</strong> With a monotone verifier, a correct proposal also certifies every superset of it. Together the correct proposals certify the region</p>

\[U(W) = \bigcup_{S \in W} \uparrow S.\]

<p>The region and \(\min_{\subseteq} W\) carry the same information: the minimal proposals generate the region, and the lowest points of the region are exactly the minimal proposals. A duplicate, or a correct proposal that contains another one, does not change anything. So any group value that ignores duplicates and redundant proposals, as our objective does, can be written as a value of the region alone, \(G(W) = R(U(W))\) [5].</p>

<p><strong>Valuing the region.</strong> We value a region by its coverage measure: its probability \(R(U) = \mu(U)\) under a base measure \(\mu\)—a product measure that includes each element independently with probability \(p\). Think of \(\mu\) as drawing a random recipe, with a coin flip for each ingredient that puts it in with probability \(p\); the value of a region is the chance that this random recipe falls inside it, that is, contains one of our correct proposals. A single correct set is then worth \(\mu(\uparrow S) = p^{\lvert S\rvert}\), and this gives us both things we asked for. A smaller set has a larger up-set, so removing a useless ingredient increases the value. And because incomparable sets cover different parts of the lattice, a new alternative will also increase the value by covering what the others did not. The product form is also the only one under which removing an ingredient multiplies the value by the same factor at every size [5]. This matters because a redundant answer can be long. If a set were worth \(1/(\lvert S\rvert + 1)\) instead, removing an ingredient would increase the value of a two-ingredient set by half but that of a twenty-ingredient set by only 5%, and the credit would hardly push a long answer to shrink. Under the product measure, every ingredient removed multiplies the value by \(1/p\), however long the answer is.</p>

<p><strong>Credit by deletion.</strong> Each proposal gets the value its group loses without it,</p>

\[A_i = R(U) - R(U_{-i}) = y_i\,\mu\big(\uparrow S_i \setminus U_{-i}\big),\]

<p>where \(U_{-i}\) is the region the other proposals certify. Since the subtracted term does not depend on proposal \(i\), the policy gradient stays unbiased.</p>

<p>This credit gives back exactly the objective we started from. Because the product measure gives every recipe a positive chance (as long as \(0 &lt; p &lt; 1\)), \(A_i &gt; 0\) exactly when proposal \(i\) is correct and no other correct proposal of its group is contained in it [5]. The reason is obvious. Take a correct proposal \(i\). Its set \(S_i\) is itself a point of \(\uparrow S_i\), and another correct proposal \(T\) certifies that point only when \(T \subseteq S_i\). If there is no such \(T\), proposal \(i\) alone certifies \(S_i\), which has positive probability, so \(A_i &gt; 0\). If there is one, its up-set already covers all of \(\uparrow S_i\), and deleting proposal \(i\) does not lose anything. In other words, the proposals with positive credit are the members of \(\min_{\subseteq} W\) that appear only once in the group. One caveat: group-minimal is weaker than minimal. A redundant set still earns credit when its group happens to miss the correct sets inside it. So unlike other group policy gradient methods, where increasing the group size only reduces the variance, in our case the group size is actually a scaling factor, as group-minimal approximates the global minimality when the group size is large enough [5]. Figure 3 computes this credit on the lattice of Figure 1.</p>

<figure class="fig" data-fig="credit">
  <div class="fig__canvas"></div>
  <figcaption class="fig__caption"><b>Figure 3</b>Deletion credit \(A_i\) of each proposal in a group, with \(p = 0.7\). Click a recipe to add it to the group or remove it; hover a row to see the part of the region only that proposal certifies.</figcaption>
</figure>

<p>You may notice that \(R(U)\) is a coverage function, which is submodular, and the credit is just its marginal gain, as in submodular RL [14]. The form also looks like the counterfactual credit of COMA [15] and the leave-one-out baseline of RLOO [9, 10]. The difference is in what gets deleted, and why. COMA removes an agent’s action to ask how much that agent added to the team’s shared reward, a counterfactual that splits credit among agents. RLOO removes a rollout’s reward from the group average only to get a baseline that reduces variance; each rollout is still judged by its own reward. Here, deleting a proposal removes the part of the lattice that only it certifies. Like COMA’s, this is a counterfactual, but over the answer instead of a reward. So the credit measures how much of the answer a proposal supplies instead of how much reward it gets. In the paper, we also introduce a leave-two-out variant of this credit (l2o) that reduces its variance [5].</p>

<h2 id="anatomy">Anatomy</h2>

<p>We first watch both mechanisms on the recipe lattice of Figure 1: the two planners one step at a time, then the two policy gradients in real time.</p>

<p>Scalar value iteration starts from what stopping pays, \(s(S) - \lambda\lvert S\rvert\), and repeats</p>

\[V_{k+1}(S) = \max\Big(s(S) - \lambda\lvert S\rvert,\ \max_{e \notin S} V_k\big(S \cup \{e\}\big)\Big)\]

<p>until nothing changes; the greedy policy then walks up from the empty recipe. Minimal-witness value iteration plans over regions instead. Its state is the region \(F\) certified by the recipes committed so far, and committing a correct recipe \(S\) adds the coverage</p>

\[\Delta(S \mid F) = R(F \cup \uparrow S) - R(F) = \mu\big(\uparrow S \setminus F\big).\]

<p>Starting from \(F_0 = \varnothing\), it repeats</p>

\[F_{k+1} = F_k \cup \uparrow S_k, \qquad S_k \in \arg\max_{S:\, s(S) = 1} \Delta(S \mid F_k),\]

<p>stopping when the largest \(\Delta\) is zero. Each step then commits a new minimal recipe, and the planner is guaranteed to stop when it has recovered the entire antichain [5].</p>

<figure class="fig" data-fig="vi">
  <div class="fig__canvas"></div>
  <figcaption class="fig__caption"><b>Figure 4</b>Value iteration on the lattice of Figure 1. Top: scalar value iteration with a size penalty of 0.1 per ingredient; the numbers are state values. Bottom: minimal-witness value iteration with \(p = 0.7\); the numbers are the coverage each correct recipe would add.</figcaption>
</figure>

<figure class="fig" data-fig="pg">
  <div class="fig__canvas"></div>
  <figcaption class="fig__caption"><b>Figure 5</b>Policy-gradient training on the lattice of Figure 1, for a tabular policy over the 16 recipes and groups of 8, one update per frame; the numbers are the policy's probabilities. GRPO normalizes the 0/1 correctness rewards within the group (learning rate 0.5); MWRL uses the deletion credit with the leave-two-out baseline and \(p = 0.7\) (learning rate 20). Restart draws a new seed.</figcaption>
</figure>

<h2 id="prime-implicant-enumeration">Prime-implicant enumeration</h2>

<p>The prime implicants of a monotone Boolean formula form a ground-truth antichain that we can enumerate (much like a maze or grid world you see in other RL papers), so we can demonstrate the different return structure from many RL methods. As the figure shows, our method recovers most of the antichain.</p>

<figure class="fig" data-fig="maxsat">
  <div class="fig__canvas" data-src="/assets/data/rlvr-maxsat.json"></div>
  <figcaption class="fig__caption"><b>Figure 6</b>Rollouts of each trained policy on one monotone MaxSAT instance with 14 variables and 10 clauses of length 3, whose satisfying assignments have 19 minimal elements. Each objective is trained with the paper's settings (groups of 48, 150 updates, seed 0) and then sampled 256 times.</figcaption>
</figure>

<h2 id="ai4s">AI4S</h2>

<p>Previous prime-implicant enumeration methods cannot amortize. But for AI4S, amortization matters. We train one policy conditioned on substrate across many Suzuki–Miyaura substrate pairs to propose the minimal sets of reaction conditions that must leave a baseline protocol, and test it on pairs it never saw. The substrate pairs are real reactions from the literature, but no existing dataset gives the yields of whole condition panels across many substrate pairs: papers report only a few conditions per reaction, mostly the ones that worked. So we fill in the panels with a mechanistic simulator calibrated on measured screens, and the yields here are simulated, not measured [5].</p>

<figure class="fig" data-fig="suzuki">
  <div class="fig__canvas" data-src="/assets/data/rlvr-suzuki.json"></div>
  <figcaption class="fig__caption"><b>Figure 7</b>Held-out Suzuki–Miyaura substrate pairs, never seen in training. Each row of the switchboard is one minimal set of the 14 condition dimensions that must leave the baseline protocol (Pd(PPh<sub>3</sub>)<sub>4</sub>, Na<sub>2</sub>CO<sub>3</sub>, dioxane with 20% water, 80 °C, 4 h) together to reach 75% yield, with the values of its best reaction; the strips are the first 32 proposals of the fingerprint-conditioned and of the substrate-blind policy. Below, recall on all 205 held-out pairs, each policy sorted by its own recall; the shaded gap is what conditioning on the substrate adds. One run (seed 0) of the mechanistic benchmark: 819 training pairs, 4,800 updates.</figcaption>
</figure>

<h2 id="mechanistic-interpretability">Mechanistic interpretability</h2>

<p>Interestingly, we can see the discovery of sparse circuits in a dense LLM as a sufficiency question (“What are the minimal sparse circuits that do not affect the performance?”). We also chose this case on purpose because monotonicity does not strictly hold here, to test robustness; the guarantee then falls back to <span class="lining">1-minimality</span> [16], where no single component can be removed. For each of 10 MMLU subjects we recover the family of <span class="lining">1-minimal</span> circuits in a frozen <span class="lining">Qwen3-1.7B</span>.</p>

<figure class="fig" data-fig="circuits">
  <div class="fig__canvas" data-src="/assets/data/rlvr-circuits.json"></div>
  <figcaption class="fig__caption"><b>Figure 8</b>1-minimal circuits recovered in Qwen3-1.7B (28 layers × 16 heads, MLP blocks below) for 10 MMLU subjects. A set of heads and MLP blocks is sufficient when, with every other component resample-ablated, it retains at least 80% of the clean-to-fully-ablated KL gap; the circuits of one subject form an antichain.</figcaption>
</figure>

<h2 id="llm-post-training">LLM Post-training</h2>

<p>We post-train a <span class="lining">Qwen3-4B</span> model with MWRL and other RL methods, including MaxRL [17], on problems whose answers actually form an antichain. The task is GMAT-style data sufficiency: each problem lists 8 true statements about two hidden whole numbers \(x\) and \(y\), and the model must name a minimal set of statements that determines \(x\) and give the value of \(x\).</p>

<figure class="fig" data-fig="training">
  <div class="fig__canvas" data-src="/assets/data/rlvr-training.json"></div>
  <figcaption class="fig__caption"><b>Figure 9</b>Training curves of Qwen3-4B-Base. All seven methods share the model, the data and the budget (256 problems with 16 rollouts per step). Every point is measured on the 16 rollouts of problems the policy has not seen, and the dotted lines mark the base model.</figcaption>
</figure>

<p>On 512 held-out problems, MWRL finds about 70% of a problem’s minimal answers within 64 samples, about twice as many as any other method, including GRPO with a minimality oracle (Table 1). The size penalty and the oracle still score each proposal alone, so a repeated answer earns as much as a new one and their policies settle on a few small answers; the diversity reward compares proposals by their raw sets, so a superset of a found answer counts as new and its policy spreads over supersets.</p>

<figure class="fig fig--table">
  <div class="fig__scroll">

    <table>
      <thead>
        <tr>
          <th>Method</th>
          <th>Found by 16 samples</th>
          <th>Found by 64 samples</th>
          <th>Minimal</th>
          <th>Statements</th>
          <th>Distinct sets</th>
          <th>Correct</th>
        </tr>
      </thead>
      <tbody>
        <tr>
          <td>Qwen3-4B-Base</td>
          <td>4.9%</td>
          <td>16.2%</td>
          <td>18.9%</td>
          <td>4.39</td>
          <td>5.4</td>
          <td>9.4%</td>
        </tr>
        <tr>
          <td>GRPO</td>
          <td>0.0%</td>
          <td>0.0%</td>
          <td>0.0%</td>
          <td>8.00</td>
          <td>1.1</td>
          <td>99.6%</td>
        </tr>
        <tr>
          <td>MaxRL</td>
          <td>0.1%</td>
          <td>0.3%</td>
          <td>0.0%</td>
          <td>7.98</td>
          <td>1.4</td>
          <td>99.0%</td>
        </tr>
        <tr>
          <td>GRPO + size penalty</td>
          <td>28.0%</td>
          <td>33.2%</td>
          <td>84.6%</td>
          <td>2.36</td>
          <td>2.5</td>
          <td>95.5%</td>
        </tr>
        <tr>
          <td>GRPO + minimality oracle</td>
          <td>30.4%</td>
          <td>36.5%</td>
          <td>96.9%</td>
          <td>2.21</td>
          <td>2.2</td>
          <td>90.5%</td>
        </tr>
        <tr>
          <td>Diversity reward</td>
          <td>1.0%</td>
          <td>3.4%</td>
          <td>0.4%</td>
          <td>5.72</td>
          <td>32.5</td>
          <td>91.0%</td>
        </tr>
        <tr>
          <td>Diversity reward + size penalty</td>
          <td>8.8%</td>
          <td>22.5%</td>
          <td>4.2%</td>
          <td>4.51</td>
          <td>30.6</td>
          <td>87.0%</td>
        </tr>
        <tr>
          <td>MWRL</td>
          <td>41.4%</td>
          <td>72.0%</td>
          <td>37.3%</td>
          <td>2.90</td>
          <td>14.0</td>
          <td>63.0%</td>
        </tr>
      </tbody>
    </table>

  </div>
  <figcaption class="fig__caption"><b>Table 1</b>Held-out evaluation on 512 problems, 64 samples each, after training with the budget of Figure 9. Found by \(k\) samples: the share of a problem's minimal answers among \(k\) samples. Minimal: the share of correct answers that are minimal. Statements: per correct answer. Distinct sets: distinct correct sets per problem. Correct: the share of correct samples. The size penalty scales the reward by \(1 - 0.1\lvert S\rvert\).</figcaption>
</figure>

<p>MWRL does not seem to perform well on two columns, Correct and Minimal, which may look odd given the title. This is because both columns score single samples, while the actual question asks about the whole family, so a policy that collapses onto one correct answer naturally scores high on both. GRPO is the plainest case: it answers with all eight statements, the “include everything” answer from the opening, so it is correct 99.6% of the time and finds 0% of the minimal answers. The size penalty and the oracle do not fall back on including everything, but they also concentrate on a few small answers, only 2.2 to 2.5 distinct correct sets per problem, while MWRL spreads over about 14, some of them likely harder to get right. Meanwhile, the non-minimal samples that come with exploration are cheap to clean up: in a batch, any correct answer that contains a smaller correct one can be dropped without another call to the verifier. A minimal answer that was never sampled, however, cannot be recovered by any post-processing, so Found by \(k\) samples is the column that nothing else can make up for. This column also already pays for MWRL’s mistakes: it counts all \(k\) samples, incorrect ones included, so MWRL finds about twice as many minimal answers within 64 samples even though 37% of its samples are incorrect. MWRL’s 37.3% comes partly from exploration itself and partly from the caveat in the credit section: group-minimal is weaker than minimal, but given enough compute for a larger group (I ran each method with only 4 <span class="lining">L40</span> GPUs, where 100 steps with 16 answers per problem already take 2 to 3 days), we are likely to see a steady improvement.</p>

<h2 id="open-directions">Open directions</h2>

<p>MWRL is our first effort in this direction. Beyond those listed in Appendix I of [5], we raise several questions that may be worth a look.</p>

<table>
  <thead>
    <tr>
      <th>Area</th>
      <th>Direction</th>
      <th>The antichain, and how to test it</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>RLVR post-training</td>
      <td>Premise selection in Lean</td>
      <td>The minimal sufficient premise sets of a theorem; the prover is the verifier, and Mathlib [18] and LeanDojo [19] supply the data.</td>
    </tr>
    <tr>
      <td> </td>
      <td>Overthinking as redundant supersets</td>
      <td>Read a solution as a set of reasoning steps. Test whether rewarding correctness alone makes redundant steps grow over training, and whether relational credit removes them without losing accuracy, compared with a length penalty. This is the one I’m most fascinated by.</td>
    </tr>
    <tr>
      <td>Agents, software and safety</td>
      <td>Delta debugging with LLM agents</td>
      <td>The minimal failure-inducing inputs, one per cause of a bug [16].</td>
    </tr>
    <tr>
      <td> </td>
      <td>Minimal sufficient evidence for RAG</td>
      <td>Independent lines of support. The verifier, an NLI judgement, is not monotone in this case. This is where the existential closure of [5] applies.</td>
    </tr>
    <tr>
      <td>AI4S</td>
      <td>Minimal cut sets of metabolic networks</td>
      <td>Flux balance analysis is the verifier [20]; on small models the ground truth can be enumerated exactly.</td>
    </tr>
    <tr>
      <td> </td>
      <td>Minimal sufficient combinations of interventions</td>
      <td>Synthetic-lethal gene pairs, drug combinations, or the minimal sets of transcription factors that induce pluripotency, as Yamanaka’s factors do [21].</td>
    </tr>
    <tr>
      <td> </td>
      <td>Minimal sufficient substructures of molecules</td>
      <td>The substructures sufficient for a property, ordered by subgraph inclusion.</td>
    </tr>
    <tr>
      <td>Methods and theory</td>
      <td>Datasets whose answers are antichains</td>
      <td>Most datasets still assume that a problem has only one solution.</td>
    </tr>
    <tr>
      <td> </td>
      <td>A multi-answer RLVR benchmark</td>
      <td>Tasks whose answers form an antichain by nature (prime implicants, minimal hitting sets, minimal unsatisfiable cores, minimal test cases, even math reasoning), each with exact ground truth, to measure RLVR methods that respect the partial order.</td>
    </tr>
    <tr>
      <td> </td>
      <td>Coverage measures and hypervolume</td>
      <td>Unifying the two brings the tools of multi-objective optimization into policy learning, and carries the unbiased group credit of MWRL into multi-objective RL.</td>
    </tr>
    <tr>
      <td> </td>
      <td>Coverage memory across training</td>
      <td>Merging every witness found so far into the covered region lets the partial order define novelty over the whole of training, beyond a single group.</td>
    </tr>
    <tr>
      <td> </td>
      <td>Circuits of SAE features</td>
      <td>On public sparse autoencoders such as Gemma Scope [22], the minimal feature sets that reproduce a behaviour, one per mechanism.</td>
    </tr>
  </tbody>
</table>

<h2 id="references">References</h2>

<ol>
  <li>P. I. Frazier. <a href="https://arxiv.org/abs/1807.02811">A tutorial on Bayesian optimization</a>. <span class="lining">arXiv:1807.02811</span>, 2018.</li>
  <li>B. J. Shields, J. Stevens, J. Li, M. Parasram, F. Damani, J. I. M. Alvarado, J. M. Janey, R. P. Adams and A. G. Doyle. Bayesian reaction optimization as a tool for chemical synthesis. <em>Nature</em> 590, 89–96, 2021.</li>
  <li>W. V. Quine. The problem of simplifying truth functions. <em>The American Mathematical Monthly</em> 59(8), 521–531, 1952.</li>
  <li>E. J. McCluskey. Minimization of Boolean functions. <em>Bell System Technical Journal</em> 35(6), 1417–1444, 1956.</li>
  <li>T. Y. Tsui, Z. Ye, P. Cai, Y. Li, Y. Li and Z. Ai. <a href="https://arxiv.org/abs/2610.07226">Minimal Witness Reinforcement Learning</a>. <span class="lining">arXiv:2610.07226</span>, 2026.</li>
  <li>B. A. Davey and H. A. Priestley. <em>Introduction to Lattices and Order</em>, 2nd edition. Cambridge University Press, 2002.</li>
  <li>J. Schulman, F. Wolski, P. Dhariwal, A. Radford and O. Klimov. <a href="https://arxiv.org/abs/1707.06347">Proximal policy optimization algorithms</a>. <span class="lining">arXiv:1707.06347</span>, 2017.</li>
  <li>Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu and D. Guo. <a href="https://arxiv.org/abs/2402.03300">DeepSeekMath: Pushing the limits of mathematical reasoning in open language models</a>. <span class="lining">arXiv:2402.03300</span>, 2024.</li>
  <li>W. Kool, H. van Hoof and M. Welling. <a href="https://openreview.net/forum?id=r1lgTGL5DE">Buy 4 REINFORCE samples, get a baseline for free!</a> <em>ICLR Workshop on Deep Reinforcement Learning Meets Structured Prediction</em>, 2019.</li>
  <li>A. Ahmadian, C. Cremer, M. Gallé, M. Fadaee, J. Kreutzer, O. Pietquin, A. Üstün and S. Hooker. <a href="https://arxiv.org/abs/2402.14740">Back to basics: Revisiting REINFORCE-style optimization for learning from human feedback in LLMs</a>. <em>ACL</em>, 2024.</li>
  <li>C. Walder and D. Karkhanis. <a href="https://arxiv.org/abs/2505.15201">Pass@K policy optimization: Solving harder reinforcement learning problems</a>. <span class="lining">arXiv:2505.15201</span>, 2025.</li>
  <li>T. Haarnoja, A. Zhou, P. Abbeel and S. Levine. <a href="https://arxiv.org/abs/1801.01290">Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor</a>. <em>ICML</em>, 2018.</li>
  <li>E. Bengio, M. Jain, M. Korablyov, D. Precup and Y. Bengio. <a href="https://arxiv.org/abs/2106.04399">Flow network based generative models for non-iterative diverse candidate generation</a>. <em>NeurIPS</em>, 2021.</li>
  <li>M. Prajapat, M. Mutný, M. N. Zeilinger and A. Krause. <a href="https://arxiv.org/abs/2307.13372">Submodular reinforcement learning</a>. <em>ICLR</em>, 2024.</li>
  <li>J. Foerster, G. Farquhar, T. Afouras, N. Nardelli and S. Whiteson. <a href="https://arxiv.org/abs/1705.08926">Counterfactual multi-agent policy gradients</a>. <em>AAAI</em>, 2018.</li>
  <li>A. Zeller and R. Hildebrandt. Simplifying and isolating failure-inducing input. <em>IEEE Transactions on Software Engineering</em> 28(2), 183–200, 2002.</li>
  <li>F. Tajwar, G. Zeng, Y. Zhou, Y. Song, D. Arora, Y. Jiang, J. Schneider, R. Salakhutdinov, H. Feng and A. Zanette. <a href="https://arxiv.org/abs/2602.02710">Maximum likelihood reinforcement learning</a>. <em>ICML</em>, 2026.</li>
  <li>The mathlib Community. The Lean mathematical library. <em>Proceedings of the 9th ACM SIGPLAN International Conference on Certified Programs and Proofs</em>, 367–381, 2020.</li>
  <li>K. Yang, A. M. Swope, A. Gu, R. Chalamala, P. Song, S. Yu, S. Godil, R. Prenger and A. Anandkumar. LeanDojo: Theorem proving with retrieval-augmented language models. <em>NeurIPS</em>, 2023.</li>
  <li>J. D. Orth, I. Thiele and B. Ø. Palsson. What is flux balance analysis? <em>Nature Biotechnology</em> 28, 245–248, 2010.</li>
  <li>K. Takahashi and S. Yamanaka. Induction of pluripotent stem cells from mouse embryonic and adult fibroblast cultures by defined factors. <em>Cell</em> 126(4), 663–676, 2006.</li>
  <li>T. Lieberum, S. Rajamanoharan, A. Conmy, L. Smith, N. Sonnerat, V. Varma, J. Kramár, A. Dragan, R. Shah and N. Nanda. <a href="https://arxiv.org/abs/2408.05147">Gemma Scope: Open sparse autoencoders everywhere all at once on <span class="lining">Gemma 2</span></a>. <span class="lining">arXiv:2408.05147</span>, 2024.</li>
</ol>]]></content><author><name>Tuen-Yue Tsui</name><email>tytsui@seas.upenn.edu</email><uri>https://tytsui.com/</uri></author><summary type="html"><![CDATA[Many tasks have more than one minimal answer, none containing another. A reward that scores each rollout on its own cannot tell a new answer from a redundant copy of one it already has.]]></summary></entry></feed>