研究
我研究高维、非线性且含离散决策的动态经济模型。一方面,我把经济结构——可行域、状态转移、离散选择与条件期望——嵌入神经网络与强化学习,让网络只学习真正未知的部分;另一方面,我把 AI 主体放进经济系统,研究技术扩散、劳动替代、税收与货币涌现。我始终以可行性、Euler 与 Bellman 残差、市场出清、基准解和分布矩来检验结果,而不只看预测损失。
工作论文
Informed Actor-Critic Method for Optimal Control in Economic Dynamics.(与 Kenneth L. Judd、Bo Li、Serguei Maliar 合作)
求职论文(Job Market Paper)。
摘要:This paper applies deep reinforcement learning to high-dimensional optimal control in dynamic economic models. We introduce an informed actor–critic (iAC) framework that embeds structural equations and constraints directly into the policy optimization loop. Unlike standard model-free reinforcement learning algorithms that treat the economy as a black box, the iAC method leverages analytical economic transitions. Computational applications to benchmark consumption–savings, Krusell–Smith, and overlapping-generations economies show that iAC reduces the need for exploration noise and provides a precise, scalable, and sample-efficient framework for complex macroeconomic control problems.
[论文]
Solving Discrete-Continuous Dynamic Choice Models: A Soft-to-Hard AI Solver with an Application to Sovereign Default.(与 Bo Li、Serguei Maliar 合作)
摘要:To solve discrete-continuous dynamic choice models, the quantitative literature has long relied on effective permanent structural smoothing to handle non-differentiable kinks. While traditional methods produce robust qualitative insights, their permanent modifications can alter quantitative predictions. To solve the true, unsmoothed economy, we introduce a soft-to-hard (StH) deep reinforcement learning framework. By systematically annealing temporary smoothing to zero, StH enables gradient propagation across decision boundaries before enforcing hard-switch operators. We implement this framework in two complementary ways. The grid-based StH-Q implementation delivers high numerical precision in the low-dimensional benchmark, whereas the continuous-state, continuous-action StH-AC implementation avoids storing a full state–action table and provides a scalable route to higher-dimensional applications.
[论文]
W-exp: A Routed Post-Decision Expectation Adapter for Dynamic Economic Solvers.(与 Bo Li、Serguei Maliar 合作)
摘要:This paper proposes W-exp, an amortized expectation module that converts simulated shock information into a reusable learned conditional expectation function over post-decision states. W-exp is algorithm-agnostic and improves the expectation-construction step shared by neural and simulator-based solvers. Under matched transition-sample budgets, W-exp lowers final policy-function error relative to Monte Carlo benchmarks by 76.2% with Q-learning, 31.4% with informed actor-critic, and 26.9% with soft actor-critic in a consumption–saving model. In a two-shock Krusell–Smith economy, a solver using the learned exact-label variant attains 42.9% lower policy-function error than the same solver using direct exact integration.
[论文]
Gradient Descent of the Discrete-Time Ramsey Growth Model.(与 Bo Li、Yapeng Qian、Shenghao Zhu 合作)
摘要:This paper provides an optimization-oriented study of the discrete-time Ramsey growth model, developed as a companion to our continuous-time analysis, replacing every continuous-time device by its discrete-time counterpart. We derive the discrete Euler–Lagrange residual as a functional gradient and use gradient-based procedures to validate the Ramsey dynamics: strict concavity and global optimality from the second variation, local strong convexity from a discrete Poincaré inequality, and exponential convergence of the sequence-space gradient flow from a Lyapunov–Grönwall argument. We construct a quadratic Lyapunov function and show, from strict monotonicity of the optimal capital policy rather than by linearization, that the optimal policy never overshoots, certifying global convergence to the steady state. We then relate the resulting schemes to deep deterministic policy gradient methods, identifying exact analytical counterparts for residual relaxation, target networks, the actor–critic split, and off-policy sampling.
Gradient Descent of Ramsey Growth Model.(与 Bo Li、Yapeng Qian、Shenghao Zhu 合作)
摘要:This paper provides an optimization-oriented study of the continuous-time Ramsey growth model, connecting classical stability analysis with residual-minimization approaches for solving Hamilton–Jacobi–Bellman (HJB) equations. We revisit the planner's primal formulation through a variational perspective and validate the Ramsey dynamics numerically. We construct a Lyapunov function and show that it decreases along saddle-point stable trajectories, certifying convergence to the steady state. We then formulate a residual loss functional induced by the HJB operator and study its analytical properties, including well-posedness, regularity, and convexity in an infinite-dimensional setting, enabling a principled comparison of first-order and second-order optimization schemes. We extend the construction to deterministic and stochastic environments, discussing how diffusion terms modify the HJB residual.
缴费基数约束与弹性退休:养老保险改革的政策权衡及其组合效应(李博、陆毅、张培元)
摘要:本文构建包含内生退休选择的异质性个体世代交叠模型,量化评估缴费基数调整与弹性退休的独立效应和组合效应。模型同时刻画缴费基数上下限、缴费历史和双账户待遇规则,将缴费端的非线性约束、待遇端的历史依赖与退休选择纳入同一家庭决策。量化结果表明,缴费基数下限是主要政策边际。降低下限虽能减轻低收入阶段的缴费负担,却会降低养老金替代率并扩大基金赤字,宏观总量和福利改善有限。弹性退休通过少数当前生产率较高、历史缴费基数较低的边际个体延长劳动和缴费期、推迟待遇领取,从而提高产出并改善基金收支。两项政策并行能够保留缴费减负效果,并基本抵消降低下限造成的待遇下降和基金压力。Shapley 分解显示,缴费减负主要来自降低下限,基金压力下降、产出和消费增量主要来自弹性退休。拓展分析发现,将缴费窗口前移至青年期后,弹性退休缩小赤字的幅度由 4.65% 降至 0.58%。本文在统一框架中识别了缴费减负与退休激励的政策分工,并揭示延迟退休的财政作用取决于继续工作是否继续缴费。
AI Taxing AI: A Multi-Agent Reinforcement Learning Approach to Optimal Taxation.(与 Bo Li 合作)
摘要:Artificial intelligence (AI) is improving productivity, yet its economy-wide diffusion may reshape the labor market and the distribution of income. To address these issues, we develop a Multi-Agent Reinforcement Learning (MARL) model featuring endogenous skill accumulation. We find that while AI boosts aggregate output through rapid skill sharing, it displaces middle-skilled human workers and increases inequality. However, optimal taxation can redistribute the technological surplus to achieve Pareto improvements. To find the optimal tax schedule, we employ Dual-Loop Multi-Agent Reinforcement Learning to account for both the government's fiscal policy and agents' behavior. Furthermore, we demonstrate that a “Digital Commons” regime—combining universal access with real-time knowledge sharing—maximizes social welfare by mitigating the trade-off between efficiency and equity.
[论文]
Forecasting Crude Oil Spot Prices Using a Transformer-BiLSTM Architecture with NRBO-Based Hyperparameter Optimization.(与 Wenhui Huang 等合作)
《Sustainable Futures》返修中。
结合 Transformer、BiLSTM 与超参数优化,开展 Brent 和 WTI 原油现货价格预测。
进行中的研究
Reduced-State Deep Learning for Heterogeneous Agent Models.
以低维矩向量近似分布状态,使含总量不确定性的异质性主体模型保持可计算。
Emergence of Universal Equivalents with Deep Reinforcement Learning.
构建多智能体强化学习环境,使主体从自给自足演化至物物交换并收敛到一般等价物,从而在不依赖强制度假设的前提下研究价格发现与货币涌现。
发表论文
Asset Pricing in China's Stock Market.(与 Shuaiyu Jiang 合作)
Finance & Economics Vision, 1(1), 2022。
摘要:This paper uses traditional machine learning methods and deep neural networks based on both firm-specific characteristics and macroeconomic variables to price China's A-share stock market. We give the stochastic discount factor a flexible form and compare different models' performances. Since the Chinese government adopts various policies to maintain financial stability, we borrow the idea from generative adversarial networks to find the true SDF by selecting moment conditions that minimize return volatility.
[论文]
