Research
Planning, discourse, and the questions with more than one good answer.
Language models are increasingly asked to do things that don't have one right answer: assemble a plan under constraints, detect an argument's validity, propose ideas. My research builds and evaluates models for exactly these settings — where the interesting question is rarely “is this correct?” and usually “is this good, and by whose standard?” I care as much about the measurement as the modeling: benchmarks and metrics that stay honest when the answer space is open.
Current threads
Planning & agents
How language models and reasoning models can divide cognitive labor. LRPlan pairs them as collaborating agents in a domain-independent framework for planning under implicit and explicit constraints — state of the art on TravelPlanner and TimeArena-Static, at superior cost efficiency. I'm interested in what else heterogeneous model teams can do that homogeneous ones can't.
representative: LRPlan, Findings of EMNLP 2025 · [code]
Fallacies in discourse
Argument quality and fallacies in real discourse — what separates reasoning that's persuasive from reasoning that's sound, and whether models can tell the difference. Because these judgments are ones people genuinely contest, part of this work is building resources and evaluations that take that contestation seriously instead of averaging it away.
status: active
Open-ended generation & semantic diversity
Sample a model many times on an open question and you get many answers — but how different are they really, beneath the phrasing? I work on evaluating and improving the diversity of ideas models produce, so that “generate ten ideas” yields ten ideas rather than one idea worded ten ways.
status: active