Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements Paper • 2607.28661 • Published 17 days ago • 13
Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning Paper • 2607.28478 • Published 9 days ago • 5
SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space Paper • 2608.01397 • Published 6 days ago • 9
Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs Paper • 2607.27951 • Published 9 days ago • 7
Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Paper • 2607.27888 • Published 9 days ago • 8
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities? Paper • 2608.03874 • Published 4 days ago • 13
ExplainBench: Evaluating Code Explanations from Agents Paper • 2607.26451 • Published 10 days ago • 13
RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems Paper • 2607.29241 • Published 8 days ago • 11
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Paper • 2608.00155 • Published 8 days ago • 13
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents Paper • 2608.04003 • Published 4 days ago • 31
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation Paper • 2607.29209 • Published 8 days ago • 34
Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published 8 days ago • 37
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications Paper • 2607.28617 • Published 9 days ago • 36
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Paper • 2608.01837 • Published 5 days ago • 38
Meshy T2: Fast Native Mesh Generation with Flow Matching Paper • 2607.28675 • Published 11 days ago • 53
InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis Paper • 2608.02437 • Published 5 days ago • 61