Accepted Papers
Accepted workshop papers are listed below. Paper titles link to their OpenReview pages.
Algorithmic Blindness in Large Language Models: A Calibration Study of Performance PredictionPosterBoNVoyage: Learning Better Rewards without RankingPosterCharacterizing Learning-Resistant PreferencesPosterClayBuddy: A Framework, Evaluation, & Mitigation of Coding Agent FailuresPosterCLIO-BENCH: Beyond Accuracy in AI Historical JudgmentPosterConfidence as a Compass: Regime-Dependent Dual Steering for Safety–Utility Tradeoffs in Language ModelsPosterControlling for Feature Frequency in Sparse Autoencoder Steering of Mathematical ReasoningPosterDemonstrations, CoT, and Prompting: A Theoretical Analysis of In-Context LearningPosterDemystifying Classifier-Free Guidance for Auto-Regressive Image GenerationOralDimension Is Free, Horizon Is Not: When Evolution Strategies Rival Policy Gradients for LLM Fine-TuningPosterDiscretizing Reward ModelsPosterDo as I Say, Not as I Do: Instruction-Induction Conflict in LLMsPosterDon’t Blink: Evidence Collapse during Multimodal ReasoningPosterDr. Post-Training: A Data Regularization Perspective on LLM Post-TrainingPosterEconomy of Minds: Emerging Multi-Agent Intelligence with Economic InteractionsPosterEmergent Misalignment as a Distributional Phase Transition in Rank-1 LoRA Fine-TuningOralEMO: Pretraining Mixture of Experts for Emergent ModularityPosterEncoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's ExpertisePosterEVALUATING THE ROBUSTNESS OF FACT-CHECKING LLMS ON MULTILINGUAL EVIDENCEPosterFinding the Minimal Parameter Budget for Implicit Reasoning: A Data Complexity Driven Scaling Law for Language ModelsPosterFrom Refusal to Fabrication: An Affective Cause of Hallucination That Bypasses the Honesty GatePosterGOPO: Policy optimization using ranked rewardsPosterH-Probes: Extracting Hierarchical Structures From Latent Representations of Language ModelsPosterHow Faithful Is Trajectory-Based Data Attribution? Error Sources, Remedies, and Practical GuidelinePosterHow Language Models Choose Sides: Internal Representations of Instruction HierarchyPosterHow Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model CircuitsOralINFUSER: Influence-Guided Self-Evolution Improves ReasoningPosterInterpreting Attention Sink Neurons in Large Language ModelsPosterLanguage model suffer from a curse of ambiguityPosterMany Preferences, Few Policies: Towards Scalable Language Model PersonalizationPosterMatching the Distribution, Not the Mean: Learning Human Rating Distributions with VLMsPosterMeasuring Representation Reorganization from Language Models to Vision-Language ModelsPosterMMIB: A Mechanistic Interpretability Benchmark for Vision-Language ModelsPosterMode-seeking in on-policy distillation of language modelsPosterMultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMsPosterNatively Unlearnable Large Language ModelsPosterNeuron Populations Exhibit Divergent Selectivity with ScalePosterOn-Policy Self-Distillation with Sampled Demonstrations Reduces Output DiversityPosterPeer-Predictive Self-Training for Language Model ReasoningPosterPlaying Psychic: Using Thought Trees to Predict Reasoning Models Accuracy on Coding TasksPosterPredicting Neural Scaling Laws without Training: A Data Manifold OraclePosterRepresentation Steering for Emotions in Diffusion Language ModelsPosterRethinking On-Policy Self-Distillation for Thinking ModelsPosterRoundtable Policy: Scaling Multi-Agent Reasoning with Confidence-Weighted-ConsensusPosterShorter Context, Fewer Negatives: How Information Reduction Biases LLM Scientific ReasoningPosterSingle Canonical Prompts Underestimate LLM Safety's Surface-Form SensitivityPosterSingle-cell foundation models obey data scaling laws on the optimized lossPosterSpend Less, Fit Better: Budget-Efficient Scaling Law Fitting via Active Experiment SelectionOralStochastic Collapse: Video World Models Fail to Reproduce Outcome FrequenciesPosterStructure Before Collapse: Transient semantic geometry in next-token predictionPosterTaming Polysemanticity in LLMs: Theory-Grounded Feature Recovery via Sparse AutoencodersPosterTemporal Conflict Resolution in Video Large Language ModelsPosterThe Future of Facts: Tracing the Factual Generation-Verification GapOralToken GeometryPosterTowards Physical Intuitions for Alignment Dynamics: A Case Study With Randomness CrystallizationPosterTruncation Is Not the Bottleneck: Reliable Calibration Conclusions from Truncated API Log-ProbabilitiesPosterTwistBench: Benchmarking Transformational Creativity in LLMs via Literary Plot TwistsOralWeight Decay Improves Language Model PlasticityPosterWhat AstroPT knows about galaxies, and what that can teach us about LLMsPosterWhat Black-Box Benchmarks Miss: Open-Ended Generation Under-Measures Memorization in Foundation ModelsPosterWhen Do Looped Transformers Need More Loops? Calibrated Measurement of External Stop SignalsPosterWhere did the ambiguity go? Examining how multimodal models interpret polysemous wordsOralWhere Does Social Reasoning Come From? Capability Provenance in Language ModelsPosterWhy GRPO Needs Normalization: A Local-Curvature Perspective on Adaptive GradientsPosterYou Only Judge Once: Multi-response Reward Modeling in a Single Forward PassPoster