diff --git a/.translate/state/bounded_rationality.md.yml b/.translate/state/bounded_rationality.md.yml
new file mode 100644
index 0000000..65d5921
--- /dev/null
+++ b/.translate/state/bounded_rationality.md.yml
@@ -0,0 +1,6 @@
+source-sha: 3fc4e2ab19683daf8b9ff600a81681a2430dfecd
+synced-at: "2026-07-26"
+model: claude-sonnet-5
+mode: NEW
+section-count: 6
+tool-version: 0.24.0
diff --git a/.translate/state/exchange_rate_learning.md.yml b/.translate/state/exchange_rate_learning.md.yml
new file mode 100644
index 0000000..d47bad4
--- /dev/null
+++ b/.translate/state/exchange_rate_learning.md.yml
@@ -0,0 +1,6 @@
+source-sha: 3fc4e2ab19683daf8b9ff600a81681a2430dfecd
+synced-at: "2026-07-26"
+model: claude-sonnet-5
+mode: NEW
+section-count: 7
+tool-version: 0.24.0
diff --git a/.translate/state/genetic_classifier.md.yml b/.translate/state/genetic_classifier.md.yml
new file mode 100644
index 0000000..d3d3c0a
--- /dev/null
+++ b/.translate/state/genetic_classifier.md.yml
@@ -0,0 +1,6 @@
+source-sha: 3fc4e2ab19683daf8b9ff600a81681a2430dfecd
+synced-at: "2026-07-26"
+model: claude-sonnet-5
+mode: NEW
+section-count: 8
+tool-version: 0.24.0
diff --git a/.translate/state/marimon_mcgrattan_sargent.md.yml b/.translate/state/marimon_mcgrattan_sargent.md.yml
new file mode 100644
index 0000000..e853e06
--- /dev/null
+++ b/.translate/state/marimon_mcgrattan_sargent.md.yml
@@ -0,0 +1,6 @@
+source-sha: 3fc4e2ab19683daf8b9ff600a81681a2430dfecd
+synced-at: "2026-07-26"
+model: claude-sonnet-5
+mode: NEW
+section-count: 13
+tool-version: 0.24.0
diff --git a/.translate/state/phillips_adaptive.md.yml b/.translate/state/phillips_adaptive.md.yml
index 4e4bbb3..d06c9aa 100644
--- a/.translate/state/phillips_adaptive.md.yml
+++ b/.translate/state/phillips_adaptive.md.yml
@@ -1,6 +1,6 @@
-source-sha: 370a58fa8dd0bcac220f17a59afbfb2716d04e1e
-synced-at: "2026-07-23"
+source-sha: 3fc4e2ab19683daf8b9ff600a81681a2430dfecd
+synced-at: "2026-07-26"
model: claude-sonnet-5
-mode: NEW
+mode: UPDATE
section-count: 8
-tool-version: 0.23.0
+tool-version: 0.24.0
diff --git a/.translate/state/phillips_credibility.md.yml b/.translate/state/phillips_credibility.md.yml
index 4e4bbb3..d06c9aa 100644
--- a/.translate/state/phillips_credibility.md.yml
+++ b/.translate/state/phillips_credibility.md.yml
@@ -1,6 +1,6 @@
-source-sha: 370a58fa8dd0bcac220f17a59afbfb2716d04e1e
-synced-at: "2026-07-23"
+source-sha: 3fc4e2ab19683daf8b9ff600a81681a2430dfecd
+synced-at: "2026-07-26"
model: claude-sonnet-5
-mode: NEW
+mode: UPDATE
section-count: 8
-tool-version: 0.23.0
+tool-version: 0.24.0
diff --git a/.translate/state/phillips_credible_policies.md.yml b/.translate/state/phillips_credible_policies.md.yml
new file mode 100644
index 0000000..3e41402
--- /dev/null
+++ b/.translate/state/phillips_credible_policies.md.yml
@@ -0,0 +1,6 @@
+source-sha: 3fc4e2ab19683daf8b9ff600a81681a2430dfecd
+synced-at: "2026-07-26"
+model: claude-sonnet-5
+mode: NEW
+section-count: 10
+tool-version: 0.24.0
diff --git a/.translate/state/phillips_drifts_volatilities.md.yml b/.translate/state/phillips_drifts_volatilities.md.yml
index e9254c1..c7b37a2 100644
--- a/.translate/state/phillips_drifts_volatilities.md.yml
+++ b/.translate/state/phillips_drifts_volatilities.md.yml
@@ -1,6 +1,6 @@
-source-sha: 370a58fa8dd0bcac220f17a59afbfb2716d04e1e
-synced-at: "2026-07-23"
+source-sha: 3fc4e2ab19683daf8b9ff600a81681a2430dfecd
+synced-at: "2026-07-26"
model: claude-sonnet-5
-mode: NEW
+mode: UPDATE
section-count: 10
-tool-version: 0.23.0
+tool-version: 0.24.0
diff --git a/.translate/state/phillips_escaping_nash.md.yml b/.translate/state/phillips_escaping_nash.md.yml
index 5d42b93..8677c2f 100644
--- a/.translate/state/phillips_escaping_nash.md.yml
+++ b/.translate/state/phillips_escaping_nash.md.yml
@@ -1,6 +1,6 @@
-source-sha: 370a58fa8dd0bcac220f17a59afbfb2716d04e1e
-synced-at: "2026-07-23"
+source-sha: 3fc4e2ab19683daf8b9ff600a81681a2430dfecd
+synced-at: "2026-07-26"
model: claude-sonnet-5
-mode: NEW
+mode: UPDATE
section-count: 12
-tool-version: 0.23.0
+tool-version: 0.24.0
diff --git a/.translate/state/phillips_learning.md.yml b/.translate/state/phillips_learning.md.yml
index e9b1834..74da96c 100644
--- a/.translate/state/phillips_learning.md.yml
+++ b/.translate/state/phillips_learning.md.yml
@@ -1,6 +1,6 @@
-source-sha: 370a58fa8dd0bcac220f17a59afbfb2716d04e1e
-synced-at: "2026-07-23"
+source-sha: 3fc4e2ab19683daf8b9ff600a81681a2430dfecd
+synced-at: "2026-07-26"
model: claude-sonnet-5
-mode: NEW
+mode: UPDATE
section-count: 11
-tool-version: 0.23.0
+tool-version: 0.24.0
diff --git a/.translate/state/phillips_lost_conquest.md.yml b/.translate/state/phillips_lost_conquest.md.yml
index 2242d8f..163d09d 100644
--- a/.translate/state/phillips_lost_conquest.md.yml
+++ b/.translate/state/phillips_lost_conquest.md.yml
@@ -1,6 +1,6 @@
-source-sha: 370a58fa8dd0bcac220f17a59afbfb2716d04e1e
-synced-at: "2026-07-23"
+source-sha: 3fc4e2ab19683daf8b9ff600a81681a2430dfecd
+synced-at: "2026-07-26"
model: claude-sonnet-5
-mode: NEW
+mode: UPDATE
section-count: 7
-tool-version: 0.23.0
+tool-version: 0.24.0
diff --git a/.translate/state/phillips_misspecified.md.yml b/.translate/state/phillips_misspecified.md.yml
index 54c8555..40356d9 100644
--- a/.translate/state/phillips_misspecified.md.yml
+++ b/.translate/state/phillips_misspecified.md.yml
@@ -1,6 +1,6 @@
-source-sha: 370a58fa8dd0bcac220f17a59afbfb2716d04e1e
-synced-at: "2026-07-23"
+source-sha: 3fc4e2ab19683daf8b9ff600a81681a2430dfecd
+synced-at: "2026-07-26"
model: claude-sonnet-5
-mode: NEW
+mode: UPDATE
section-count: 6
-tool-version: 0.23.0
+tool-version: 0.24.0
diff --git a/.translate/state/phillips_priors.md.yml b/.translate/state/phillips_priors.md.yml
index e9254c1..c7b37a2 100644
--- a/.translate/state/phillips_priors.md.yml
+++ b/.translate/state/phillips_priors.md.yml
@@ -1,6 +1,6 @@
-source-sha: 370a58fa8dd0bcac220f17a59afbfb2716d04e1e
-synced-at: "2026-07-23"
+source-sha: 3fc4e2ab19683daf8b9ff600a81681a2430dfecd
+synced-at: "2026-07-26"
model: claude-sonnet-5
-mode: NEW
+mode: UPDATE
section-count: 10
-tool-version: 0.23.0
+tool-version: 0.24.0
diff --git a/.translate/state/phillips_self_confirming.md.yml b/.translate/state/phillips_self_confirming.md.yml
index 2242d8f..163d09d 100644
--- a/.translate/state/phillips_self_confirming.md.yml
+++ b/.translate/state/phillips_self_confirming.md.yml
@@ -1,6 +1,6 @@
-source-sha: 370a58fa8dd0bcac220f17a59afbfb2716d04e1e
-synced-at: "2026-07-23"
+source-sha: 3fc4e2ab19683daf8b9ff600a81681a2430dfecd
+synced-at: "2026-07-26"
model: claude-sonnet-5
-mode: NEW
+mode: UPDATE
section-count: 7
-tool-version: 0.23.0
+tool-version: 0.24.0
diff --git a/.translate/state/phillips_two_stories.md.yml b/.translate/state/phillips_two_stories.md.yml
index 4e4bbb3..d06c9aa 100644
--- a/.translate/state/phillips_two_stories.md.yml
+++ b/.translate/state/phillips_two_stories.md.yml
@@ -1,6 +1,6 @@
-source-sha: 370a58fa8dd0bcac220f17a59afbfb2716d04e1e
-synced-at: "2026-07-23"
+source-sha: 3fc4e2ab19683daf8b9ff600a81681a2430dfecd
+synced-at: "2026-07-26"
model: claude-sonnet-5
-mode: NEW
+mode: UPDATE
section-count: 8
-tool-version: 0.23.0
+tool-version: 0.24.0
diff --git a/.translate/state/prospects_bounded_rationality.md.yml b/.translate/state/prospects_bounded_rationality.md.yml
new file mode 100644
index 0000000..d47bad4
--- /dev/null
+++ b/.translate/state/prospects_bounded_rationality.md.yml
@@ -0,0 +1,6 @@
+source-sha: 3fc4e2ab19683daf8b9ff600a81681a2430dfecd
+synced-at: "2026-07-26"
+model: claude-sonnet-5
+mode: NEW
+section-count: 7
+tool-version: 0.24.0
diff --git a/lectures/_static/quant-econ.bib b/lectures/_static/quant-econ.bib
index 6c7f069..2b65618 100644
--- a/lectures/_static/quant-econ.bib
+++ b/lectures/_static/quant-econ.bib
@@ -2182,6 +2182,26 @@ @article{PhelanStacchetti2001
month = {November}
}
+@article{BarroGordon1983,
+ author = {Barro, Robert J. and Gordon, David B.},
+ title = {Rules, Discretion and Reputation in a Model of Monetary Policy},
+ journal = {Journal of Monetary Economics},
+ volume = {12},
+ number = {1},
+ pages = {101--121},
+ year = {1983}
+}
+
+@article{SargentVelde1995,
+ author = {Sargent, Thomas J. and Velde, Fran\c{c}ois R.},
+ title = {Macroeconomic Features of the French Revolution},
+ journal = {Journal of Political Economy},
+ volume = {103},
+ number = {3},
+ pages = {474--518},
+ year = {1995}
+}
+
@article{APS1990,
title = {Toward a Theory of Discounted Repeated Games with Imperfect Monitoring},
author = {Abreu, Dilip and David Pearce and Ennio Stacchetti},
@@ -4104,21 +4124,50 @@ @article{Campbell1987
pages = {1249--1273},
year = {1987}
}
-% ---------------------------------------------------------------------------
-% Entries synced from QuantEcon/lecture-python.myst lectures/_static/quant-econ.bib
-% on 2026-07-24. The sync action does not carry shared assets across
-% (QuantEcon/action-translation#117), so this bibliography drifts from upstream
-% until backfilled by hand.
-% ---------------------------------------------------------------------------
-@article{Andrews1993,
- author = {Andrews, Donald W. K.},
- title = {Tests for Parameter Instability and Structural Change with Unknown Change Point},
- journal = {Econometrica},
- volume = {61},
- number = {4},
- pages = {821--856},
- year = {1993}
+@book{Sargent1999,
+ author = {Sargent, Thomas J.},
+ title = {The Conquest of American Inflation},
+ publisher = {Princeton University Press},
+ address = {Princeton, New Jersey},
+ year = {1999}
+}
+
+@article{Phelps1967,
+ author = {Phelps, Edmund S.},
+ title = {Phillips Curves, Expectations of Inflation and Optimal Unemployment over Time},
+ journal = {Economica},
+ volume = {34},
+ number = {135},
+ pages = {254--281},
+ year = {1967}
+}
+
+@article{KingWatson1994,
+ author = {King, Robert G. and Watson, Mark W.},
+ title = {The Post-War {U.S.} Phillips Curve: A Revisionist Econometric History},
+ journal = {Carnegie-Rochester Conference Series on Public Policy},
+ volume = {41},
+ pages = {157--219},
+ year = {1994}
+}
+
+@book{KushnerClark1978,
+ author = {Kushner, Harold J. and Clark, Dean S.},
+ title = {Stochastic Approximation Methods for Constrained and Unconstrained Systems},
+ publisher = {Springer-Verlag},
+ address = {New York},
+ year = {1978}
+}
+
+@article{SamuelsonSolow1960,
+ author = {Samuelson, Paul A. and Solow, Robert M.},
+ title = {Analytical Aspects of Anti-Inflation Policy},
+ journal = {American Economic Review},
+ volume = {50},
+ number = {2},
+ pages = {177--194},
+ year = {1960}
}
@article{BaxterKing1999,
@@ -4131,80 +4180,93 @@ @article{BaxterKing1999
year = {1999}
}
-@book{BenvenisteMetivierPriouret1990,
- author = {Benveniste, Albert and M\'etivier, Michel and Priouret, Pierre},
- title = {Adaptive Algorithms and Stochastic Approximations},
- publisher = {Springer-Verlag},
- address = {Berlin},
- year = {1990}
+@article{Sims1988,
+ author = {Sims, Christopher A.},
+ title = {Projecting Policy Effects with Statistical Models},
+ journal = {Revista de An\'alisis Econ\'omico},
+ volume = {3},
+ number = {1},
+ pages = {3--20},
+ year = {1988}
}
-@book{Bernanke2022,
- author = {Bernanke, Ben S.},
- title = {21st Century Monetary Policy: The {Federal} {Reserve} from the Great Inflation to {COVID-19}},
- publisher = {W. W. Norton and Company},
- address = {New York},
- year = {2022}
+@phdthesis{Chung1990,
+ author = {Chung, Heetaik},
+ title = {Did Policy Makers Really Believe in the Phillips Curve? An Econometric Test},
+ school = {University of Minnesota},
+ year = {1990}
}
-@article{BernankeMihov1998,
- author = {Bernanke, Ben S. and Mihov, Ilian},
- title = {Measuring Monetary Policy},
- journal = {Quarterly Journal of Economics},
- volume = {113},
- number = {3},
- pages = {869--902},
- year = {1998}
+@article{CooleyPrescott1973,
+ author = {Cooley, Thomas F. and Prescott, Edward C.},
+ title = {An Adaptive Regression Model},
+ journal = {International Economic Review},
+ volume = {14},
+ number = {2},
+ pages = {364--371},
+ year = {1973}
}
-@article{BrockHommes1997,
- author = {Brock, William A. and Hommes, Cars H.},
- title = {A Rational Route to Randomness},
+@article{HansenSargent1993,
+ author = {Hansen, Lars Peter and Sargent, Thomas J.},
+ title = {Seasonality and Approximation Errors in Rational Expectations Models},
+ journal = {Journal of Econometrics},
+ volume = {55},
+ number = {1--2},
+ pages = {21--55},
+ year = {1993}
+}
+
+@article{Granger1966,
+ author = {Granger, C. W. J.},
+ title = {The Typical Spectral Shape of an Economic Variable},
journal = {Econometrica},
- volume = {65},
- number = {5},
- pages = {1059--1095},
- year = {1997}
+ volume = {34},
+ number = {1},
+ pages = {150--161},
+ year = {1966}
}
-@article{Bullard1994,
- author = {Bullard, James},
- title = {Learning Equilibria},
- journal = {Journal of Economic Theory},
- volume = {64},
- number = {2},
- pages = {468--485},
- year = {1994}
+@incollection{Kreps1998,
+ author = {Kreps, David M.},
+ title = {Anticipated Utility and Dynamic Choice},
+ booktitle = {Frontiers of Research in Economic Theory: The Nancy L. Schwartz Memorial Lectures, 1983--1997},
+ editor = {Jacobs, Donald P. and Kalai, Ehud and Kamien, Morton I.},
+ publisher = {Cambridge University Press},
+ address = {Cambridge},
+ pages = {242--274},
+ year = {1998}
}
-@article{CarboniEllison2009,
- author = {Carboni, Giacomo and Ellison, Martin},
- title = {The Great Inflation and the {Greenbook}},
- journal = {Journal of Monetary Economics},
- volume = {56},
- number = {6},
- pages = {831--841},
- year = {2009}
+@incollection{Lucas1972,
+ author = {Lucas, Robert E., Jr.},
+ title = {Econometric Testing of the Natural Rate Hypothesis},
+ booktitle = {The Econometrics of Price Determination},
+ editor = {Eckstein, Otto},
+ publisher = {Board of Governors of the Federal Reserve System},
+ address = {Washington, D.C.},
+ pages = {50--59},
+ year = {1972}
}
-@article{CarterKohn1994,
- author = {Carter, C. K. and Kohn, R.},
- title = {On {Gibbs} Sampling for State Space Models},
- journal = {Biometrika},
- volume = {81},
- number = {3},
- pages = {541--553},
- year = {1994}
+@incollection{Solow1968,
+ author = {Solow, Robert M.},
+ title = {Recent Controversy on the Theory of Inflation: An Eclectic View},
+ booktitle = {Proceedings of a Symposium on Inflation: Its Causes, Consequences, and Control},
+ editor = {Rousseaus, Stephen W.},
+ publisher = {New York University},
+ address = {New York},
+ year = {1968}
}
-@article{ChoMatsui1995,
- author = {Cho, In-Koo and Matsui, Akihiko},
- title = {Induction and the {Ramsey} Policy},
- journal = {Journal of Economic Dynamics and Control},
- volume = {19},
- number = {5--7},
- pages = {1113--1140},
- year = {1995}
+@incollection{Tobin1968,
+ author = {Tobin, James},
+ title = {Discussion},
+ booktitle = {Proceedings of a Symposium on Inflation: Its Causes, Consequences, and Control},
+ editor = {Rousseaus, Stephen W.},
+ publisher = {New York University},
+ address = {New York},
+ year = {1968}
}
@article{ChoWilliamsSargent2002,
@@ -4217,87 +4279,70 @@ @article{ChoWilliamsSargent2002
year = {2002}
}
-@phdthesis{Chung1990,
- author = {Chung, Heetaik},
- title = {Did Policy Makers Really Believe in the Phillips Curve? An Econometric Test},
- school = {University of Minnesota},
- year = {1990}
-}
-
-@article{Clarida2000,
- author = {Clarida, Richard and Gal\'{i}, Jordi and Gertler, Mark},
- title = {Monetary Policy Rules and Macroeconomic Stability: Evidence and Some Theory},
- journal = {Quarterly Journal of Economics},
- volume = {115},
- number = {1},
- pages = {147--180},
- year = {2000}
-}
-
-@article{CogleySargent2001,
- author = {Cogley, Timothy and Sargent, Thomas J.},
- title = {Evolving Post-World War {II} {U.S.} Inflation Dynamics},
- journal = {NBER Macroeconomics Annual},
- volume = {16},
- pages = {331--373},
- year = {2001}
-}
-
-@article{CogleySargent2005,
- author = {Cogley, Timothy and Sargent, Thomas J.},
- title = {Drifts and Volatilities: Monetary Policies and Outcomes in the Post {WWII} {US}},
- journal = {Review of Economic Dynamics},
- volume = {8},
- number = {2},
- pages = {262--302},
- year = {2005}
-}
-
-@article{CogleySargentConquest2005,
- author = {Cogley, Timothy and Sargent, Thomas J.},
- title = {The Conquest of {US} Inflation: Learning and Robustness to Model Uncertainty},
+@article{SargentWilliams2005,
+ author = {Sargent, Thomas J. and Williams, Noah},
+ title = {Impacts of Priors on Convergence and Escapes from {Nash} Inflation},
journal = {Review of Economic Dynamics},
volume = {8},
number = {2},
- pages = {528--563},
+ pages = {360--391},
year = {2005}
}
-@article{CooleyPrescott1973,
- author = {Cooley, Thomas F. and Prescott, Edward C.},
- title = {An Adaptive Regression Model},
+@article{Kasa2004,
+ author = {Kasa, Kenneth},
+ title = {Learning, Large Deviations, and Recurrent Currency Crises},
journal = {International Economic Review},
- volume = {14},
- number = {2},
- pages = {364--371},
- year = {1973}
+ volume = {45},
+ number = {1},
+ pages = {141--173},
+ year = {2004}
}
-@incollection{DeLong1997,
- author = {DeLong, J. Bradford},
- title = {America's Peacetime Inflation: The 1970s},
- booktitle = {Reducing Inflation: Motivation and Strategy},
- editor = {Romer, Christina D. and Romer, David H.},
- publisher = {University of Chicago Press},
- pages = {247--280},
- year = {1997}
+@article{RobbinsMonro1951,
+ author = {Robbins, Herbert and Monro, Sutton},
+ title = {A Stochastic Approximation Method},
+ journal = {Annals of Mathematical Statistics},
+ volume = {22},
+ number = {3},
+ pages = {400--407},
+ year = {1951}
}
-@book{DemboZeitouni1998,
- author = {Dembo, Amir and Zeitouni, Ofer},
- title = {Large Deviations Techniques and Applications},
+@article{KieferWolfowitz1952,
+ author = {Kiefer, Jack and Wolfowitz, Jacob},
+ title = {Stochastic Estimation of the Maximum of a Regression Function},
+ journal = {Annals of Mathematical Statistics},
+ volume = {23},
+ number = {3},
+ pages = {462--466},
+ year = {1952}
+}
+
+@book{BenvenisteMetivierPriouret1990,
+ author = {Benveniste, Albert and M\'etivier, Michel and Priouret, Pierre},
+ title = {Adaptive Algorithms and Stochastic Approximations},
+ publisher = {Springer-Verlag},
+ address = {Berlin},
+ year = {1990}
+}
+
+@book{KushnerYin2003,
+ author = {Kushner, Harold J. and Yin, G. George},
+ title = {Stochastic Approximation and Recursive Algorithms and Applications},
edition = {2nd},
publisher = {Springer-Verlag},
address = {New York},
- year = {1998}
+ year = {2003}
}
-@book{DupuisEllis1997,
- author = {Dupuis, Paul and Ellis, Richard S.},
- title = {A Weak Convergence Approach to the Theory of Large Deviations},
- publisher = {John Wiley and Sons},
+@book{FreidlinWentzell1998,
+ author = {Freidlin, Mark I. and Wentzell, Alexander D.},
+ title = {Random Perturbations of Dynamical Systems},
+ edition = {2nd},
+ publisher = {Springer-Verlag},
address = {New York},
- year = {1997}
+ year = {1998}
}
@article{DupuisKushner1987,
@@ -4320,43 +4365,64 @@ @article{DupuisKushner1989
year = {1989}
}
-@article{EllisonYates2007,
- author = {Ellison, Martin and Yates, Tony},
- title = {Escaping Volatile Inflation},
- journal = {Journal of Money, Credit and Banking},
- volume = {39},
- number = {4},
- pages = {981--993},
- year = {2007}
+@book{DupuisEllis1997,
+ author = {Dupuis, Paul and Ellis, Richard S.},
+ title = {A Weak Convergence Approach to the Theory of Large Deviations},
+ publisher = {John Wiley and Sons},
+ address = {New York},
+ year = {1997}
}
-@book{FreidlinWentzell1998,
- author = {Freidlin, Mark I. and Wentzell, Alexander D.},
- title = {Random Perturbations of Dynamical Systems},
- edition = {2nd},
- publisher = {Springer-Verlag},
- address = {New York},
- year = {1998}
+@article{Woodford1990,
+ author = {Woodford, Michael},
+ title = {Learning to Believe in Sunspots},
+ journal = {Econometrica},
+ volume = {58},
+ number = {2},
+ pages = {277--307},
+ year = {1990}
}
-@book{Friedman1957,
- author = {Friedman, Milton},
- title = {A Theory of the Consumption Function},
- publisher = {Princeton University Press},
- address = {Princeton, New Jersey},
- year = {1957}
+@article{BrockHommes1997,
+ author = {Brock, William A. and Hommes, Cars H.},
+ title = {A Rational Route to Randomness},
+ journal = {Econometrica},
+ volume = {65},
+ number = {5},
+ pages = {1059--1095},
+ year = {1997}
}
-@article{FudenbergLevine1993,
- author = {Fudenberg, Drew and Levine, David K.},
- title = {Self-Confirming Equilibrium},
+@article{KandoriMailathRob1993,
+ author = {Kandori, Michihiro and Mailath, George J. and Rob, Rafael},
+ title = {Learning, Mutation, and Long Run Equilibria in Games},
journal = {Econometrica},
volume = {61},
- number = {3},
- pages = {523--545},
+ number = {1},
+ pages = {29--56},
year = {1993}
}
+@article{Williams2019,
+ author = {Williams, Noah},
+ title = {Escape Dynamics in Learning Models},
+ journal = {Review of Economic Studies},
+ volume = {86},
+ number = {2},
+ pages = {882--912},
+ year = {2019}
+}
+
+@article{Bullard1994,
+ author = {Bullard, James},
+ title = {Learning Equilibria},
+ journal = {Journal of Economic Theory},
+ volume = {64},
+ number = {2},
+ pages = {468--485},
+ year = {1994}
+}
+
@book{FudenbergLevine1998,
author = {Fudenberg, Drew and Levine, David K.},
title = {The Theory of Learning in Games},
@@ -4365,24 +4431,43 @@ @book{FudenbergLevine1998
year = {1998}
}
-@article{Granger1966,
- author = {Granger, C. W. J.},
- title = {The Typical Spectral Shape of an Economic Variable},
- journal = {Econometrica},
- volume = {34},
- number = {1},
- pages = {150--161},
- year = {1966}
+@article{SargentWilliamsZha2006,
+ author = {Sargent, Thomas J. and Williams, Noah and Zha, Tao},
+ title = {Shocks and Government Beliefs: The Rise and Fall of American Inflation},
+ journal = {American Economic Review},
+ volume = {96},
+ number = {4},
+ pages = {1193--1224},
+ year = {2006}
}
-@article{HansenSargent1993,
- author = {Hansen, Lars Peter and Sargent, Thomas J.},
- title = {Seasonality and Approximation Errors in Rational Expectations Models},
- journal = {Journal of Econometrics},
- volume = {55},
- number = {1--2},
- pages = {21--55},
- year = {1993}
+@article{CogleySargent2005,
+ author = {Cogley, Timothy and Sargent, Thomas J.},
+ title = {Drifts and Volatilities: Monetary Policies and Outcomes in the Post {WWII} {US}},
+ journal = {Review of Economic Dynamics},
+ volume = {8},
+ number = {2},
+ pages = {262--302},
+ year = {2005}
+}
+
+@article{CogleySargent2001,
+ author = {Cogley, Timothy and Sargent, Thomas J.},
+ title = {Evolving Post-World War {II} {U.S.} Inflation Dynamics},
+ journal = {NBER Macroeconomics Annual},
+ volume = {16},
+ pages = {331--373},
+ year = {2001}
+}
+
+@article{Primiceri2005,
+ author = {Primiceri, Giorgio E.},
+ title = {Time Varying Structural Vector Autoregressions and Monetary Policy},
+ journal = {Review of Economic Studies},
+ volume = {72},
+ number = {3},
+ pages = {821--852},
+ year = {2005}
}
@article{Jacquier1994,
@@ -4395,92 +4480,102 @@ @article{Jacquier1994
year = {1994}
}
-@article{KandoriMailathRob1993,
- author = {Kandori, Michihiro and Mailath, George J. and Rob, Rafael},
- title = {Learning, Mutation, and Long Run Equilibria in Games},
- journal = {Econometrica},
- volume = {61},
- number = {1},
- pages = {29--56},
- year = {1993}
+@article{CarterKohn1994,
+ author = {Carter, C. K. and Kohn, R.},
+ title = {On {Gibbs} Sampling for State Space Models},
+ journal = {Biometrika},
+ volume = {81},
+ number = {3},
+ pages = {541--553},
+ year = {1994}
}
-@article{Kasa2004,
- author = {Kasa, Kenneth},
- title = {Learning, Large Deviations, and Recurrent Currency Crises},
- journal = {International Economic Review},
- volume = {45},
- number = {1},
- pages = {141--173},
- year = {2004}
+@article{BernankeMihov1998,
+ author = {Bernanke, Ben S. and Mihov, Ilian},
+ title = {Measuring Monetary Policy},
+ journal = {Quarterly Journal of Economics},
+ volume = {113},
+ number = {3},
+ pages = {869--902},
+ year = {1998}
}
-@article{KieferWolfowitz1952,
- author = {Kiefer, Jack and Wolfowitz, Jacob},
- title = {Stochastic Estimation of the Maximum of a Regression Function},
- journal = {Annals of Mathematical Statistics},
- volume = {23},
- number = {3},
- pages = {462--466},
- year = {1952}
+@article{SimsZha2006,
+ author = {Sims, Christopher A. and Zha, Tao},
+ title = {Were There Regime Switches in {U.S.} Monetary Policy?},
+ journal = {American Economic Review},
+ volume = {96},
+ number = {1},
+ pages = {54--81},
+ year = {2006}
}
-@article{KimNelson1999,
- author = {Kim, Chang-Jin and Nelson, Charles R.},
- title = {Has the {U.S.} Economy Become More Stable? A {Bayesian} Approach Based on a {Markov}-Switching Model of the Business Cycle},
- journal = {Review of Economics and Statistics},
- volume = {81},
- number = {4},
- pages = {608--616},
- year = {1999}
+@article{Sims2001comment,
+ author = {Sims, Christopher A.},
+ title = {Comment on {Sargent} and {Cogley's} `Evolving Post World War {II} {US} Inflation Dynamics'},
+ journal = {NBER Macroeconomics Annual},
+ volume = {16},
+ pages = {373--379},
+ year = {2001}
}
-@article{KingWatson1994,
- author = {King, Robert G. and Watson, Mark W.},
- title = {The Post-War {U.S.} Phillips Curve: A Revisionist Econometric History},
- journal = {Carnegie-Rochester Conference Series on Public Policy},
- volume = {41},
- pages = {157--219},
- year = {1994}
+@article{Stock2001comment,
+ author = {Stock, James H.},
+ title = {Discussion of {Cogley} and {Sargent} `Evolving Post World War {II} {US} Inflation Dynamics'},
+ journal = {NBER Macroeconomics Annual},
+ volume = {16},
+ pages = {379--387},
+ year = {2001}
}
-@incollection{Kreps1998,
- author = {Kreps, David M.},
- title = {Anticipated Utility and Dynamic Choice},
- booktitle = {Frontiers of Research in Economic Theory: The Nancy L. Schwartz Memorial Lectures, 1983--1997},
- editor = {Jacobs, Donald P. and Kalai, Ehud and Kamien, Morton I.},
- publisher = {Cambridge University Press},
- address = {Cambridge},
- pages = {242--274},
- year = {1998}
+@article{Andrews1993,
+ author = {Andrews, Donald W. K.},
+ title = {Tests for Parameter Instability and Structural Change with Unknown Change Point},
+ journal = {Econometrica},
+ volume = {61},
+ number = {4},
+ pages = {821--856},
+ year = {1993}
}
-@book{KushnerClark1978,
- author = {Kushner, Harold J. and Clark, Dean S.},
- title = {Stochastic Approximation Methods for Constrained and Unconstrained Systems},
- publisher = {Springer-Verlag},
- address = {New York},
- year = {1978}
+@article{Nyblom1989,
+ author = {Nyblom, Jukka},
+ title = {Testing for the Constancy of Parameters over Time},
+ journal = {Journal of the American Statistical Association},
+ volume = {84},
+ number = {405},
+ pages = {223--230},
+ year = {1989}
}
-@book{KushnerYin2003,
- author = {Kushner, Harold J. and Yin, G. George},
- title = {Stochastic Approximation and Recursive Algorithms and Applications},
- edition = {2nd},
- publisher = {Springer-Verlag},
- address = {New York},
- year = {2003}
+@article{Whittle1953,
+ author = {Whittle, Peter},
+ title = {The Analysis of Multiple Stationary Time Series},
+ journal = {Journal of the Royal Statistical Society, Series B},
+ volume = {15},
+ number = {1},
+ pages = {125--139},
+ year = {1953}
}
-@incollection{Lucas1972,
- author = {Lucas, Robert E., Jr.},
- title = {Econometric Testing of the Natural Rate Hypothesis},
- booktitle = {The Econometrics of Price Determination},
- editor = {Eckstein, Otto},
- publisher = {Board of Governors of the Federal Reserve System},
- address = {Washington, D.C.},
- pages = {50--59},
- year = {1972}
+@article{Clarida2000,
+ author = {Clarida, Richard and Gal\'{i}, Jordi and Gertler, Mark},
+ title = {Monetary Policy Rules and Macroeconomic Stability: Evidence and Some Theory},
+ journal = {Quarterly Journal of Economics},
+ volume = {115},
+ number = {1},
+ pages = {147--180},
+ year = {2000}
+}
+
+@article{KimNelson1999,
+ author = {Kim, Chang-Jin and Nelson, Charles R.},
+ title = {Has the {U.S.} Economy Become More Stable? A {Bayesian} Approach Based on a {Markov}-Switching Model of the Business Cycle},
+ journal = {Review of Economics and Statistics},
+ volume = {81},
+ number = {4},
+ pages = {608--616},
+ year = {1999}
}
@article{McConnellPerezQuiros2000,
@@ -4493,46 +4588,24 @@ @article{McConnellPerezQuiros2000
year = {2000}
}
-@incollection{McLeayTenreyro2019,
- author = {McLeay, Michael and Tenreyro, Silvana},
- title = {Optimal Inflation and the Identification of the Phillips Curve},
- booktitle = {NBER Macroeconomics Annual 2018, Volume 33},
- editor = {Eichenbaum, Martin and Hurst, Erik and Parker, Jonathan A.},
+@incollection{DeLong1997,
+ author = {DeLong, J. Bradford},
+ title = {America's Peacetime Inflation: The 1970s},
+ booktitle = {Reducing Inflation: Motivation and Strategy},
+ editor = {Romer, Christina D. and Romer, David H.},
publisher = {University of Chicago Press},
- address = {Chicago},
- pages = {199--255},
- year = {2019}
-}
-
-@inproceedings{MurrayAdamsMacKay2010,
- author = {Murray, Iain and Adams, Ryan P. and MacKay, David J. C.},
- title = {Elliptical Slice Sampling},
- booktitle = {Proceedings of the Thirteenth International Conference on
- Artificial Intelligence and Statistics},
- series = {Proceedings of Machine Learning Research},
- volume = {9},
- pages = {541--548},
- year = {2010}
-}
-
-@article{Nyblom1989,
- author = {Nyblom, Jukka},
- title = {Testing for the Constancy of Parameters over Time},
- journal = {Journal of the American Statistical Association},
- volume = {84},
- number = {405},
- pages = {223--230},
- year = {1989}
+ pages = {247--280},
+ year = {1997}
}
-@article{Orphanides2001,
- author = {Orphanides, Athanasios},
- title = {Monetary Policy Rules Based on Real-Time Data},
- journal = {American Economic Review},
- volume = {91},
- number = {4},
- pages = {964--985},
- year = {2001}
+@incollection{Taylor1997comment,
+ author = {Taylor, John B.},
+ title = {Comment on `America's Peacetime Inflation: The 1970s'},
+ booktitle = {Reducing Inflation: Motivation and Strategy},
+ editor = {Romer, Christina D. and Romer, David H.},
+ publisher = {University of Chicago Press},
+ pages = {276--280},
+ year = {1997}
}
@book{Perko1996,
@@ -4544,54 +4617,43 @@ @book{Perko1996
year = {1996}
}
-@article{Phelps1967,
- author = {Phelps, Edmund S.},
- title = {Phillips Curves, Expectations of Inflation and Optimal Unemployment over Time},
- journal = {Economica},
- volume = {34},
- number = {135},
- pages = {254--281},
- year = {1967}
-}
-
-@article{PhelpsTaylor1977,
- author = {Phelps, Edmund S. and Taylor, John B.},
- title = {Stabilizing Powers of Monetary Policy under Rational Expectations},
- journal = {Journal of Political Economy},
- volume = {85},
- number = {1},
- pages = {163--190},
- year = {1977}
+@article{StockWatson1998,
+ author = {Stock, James H. and Watson, Mark W.},
+ title = {Median Unbiased Estimation of Coefficient Variance in a Time-Varying Parameter Model},
+ journal = {Journal of the American Statistical Association},
+ volume = {93},
+ number = {441},
+ pages = {349--358},
+ year = {1998}
}
-@article{Primiceri2005,
- author = {Primiceri, Giorgio E.},
- title = {Time Varying Structural Vector Autoregressions and Monetary Policy},
- journal = {Review of Economic Studies},
- volume = {72},
+@article{FudenbergLevine1993,
+ author = {Fudenberg, Drew and Levine, David K.},
+ title = {Self-Confirming Equilibrium},
+ journal = {Econometrica},
+ volume = {61},
number = {3},
- pages = {821--852},
- year = {2005}
+ pages = {523--545},
+ year = {1993}
}
-@article{Primiceri2006,
- author = {Primiceri, Giorgio E.},
- title = {Why Inflation Rose and Fell: Policymakers' Beliefs and {U.S.} Postwar Stabilization Policy},
- journal = {Quarterly Journal of Economics},
- volume = {121},
- number = {3},
- pages = {867--901},
- year = {2006}
+@article{ChoMatsui1995,
+ author = {Cho, In-Koo and Matsui, Akihiko},
+ title = {Induction and the {Ramsey} Policy},
+ journal = {Journal of Economic Dynamics and Control},
+ volume = {19},
+ number = {5--7},
+ pages = {1113--1140},
+ year = {1995}
}
-@article{RobbinsMonro1951,
- author = {Robbins, Herbert and Monro, Sutton},
- title = {A Stochastic Approximation Method},
- journal = {Annals of Mathematical Statistics},
- volume = {22},
- number = {3},
- pages = {400--407},
- year = {1951}
+@book{DemboZeitouni1998,
+ author = {Dembo, Amir and Zeitouni, Ofer},
+ title = {Large Deviations Techniques and Applications},
+ edition = {2nd},
+ publisher = {Springer-Verlag},
+ address = {New York},
+ year = {1998}
}
@article{Rogoff1985,
@@ -4604,22 +4666,52 @@ @article{Rogoff1985
year = {1985}
}
-@article{SamuelsonSolow1960,
- author = {Samuelson, Paul A. and Solow, Robert M.},
- title = {Analytical Aspects of Anti-Inflation Policy},
- journal = {American Economic Review},
- volume = {50},
- number = {2},
- pages = {177--194},
- year = {1960}
+@article{Sims1980,
+ author = {Sims, Christopher A.},
+ title = {Macroeconomics and Reality},
+ journal = {Econometrica},
+ volume = {48},
+ number = {1},
+ pages = {1--48},
+ year = {1980}
}
-@book{Sargent1999,
- author = {Sargent, Thomas J.},
- title = {The Conquest of American Inflation},
+@book{Friedman1957,
+ author = {Friedman, Milton},
+ title = {A Theory of the Consumption Function},
publisher = {Princeton University Press},
address = {Princeton, New Jersey},
- year = {1999}
+ year = {1957}
+}
+
+@article{EllisonYates2007,
+ author = {Ellison, Martin and Yates, Tony},
+ title = {Escaping Volatile Inflation},
+ journal = {Journal of Money, Credit and Banking},
+ volume = {39},
+ number = {4},
+ pages = {981--993},
+ year = {2007}
+}
+
+@article{CarboniEllison2009,
+ author = {Carboni, Giacomo and Ellison, Martin},
+ title = {The Great Inflation and the {Greenbook}},
+ journal = {Journal of Monetary Economics},
+ volume = {56},
+ number = {6},
+ pages = {831--841},
+ year = {2009}
+}
+
+@article{Primiceri2006,
+ author = {Primiceri, Giorgio E.},
+ title = {Why Inflation Rose and Fell: Policymakers' Beliefs and {U.S.} Postwar Stabilization Policy},
+ journal = {Quarterly Journal of Economics},
+ volume = {121},
+ number = {3},
+ pages = {867--901},
+ year = {2006}
}
@article{Sargent2008,
@@ -4632,14 +4724,14 @@ @article{Sargent2008
year = {2008}
}
-@article{SargentWilliams2005,
- author = {Sargent, Thomas J. and Williams, Noah},
- title = {Impacts of Priors on Convergence and Escapes from {Nash} Inflation},
- journal = {Review of Economic Dynamics},
- volume = {8},
- number = {2},
- pages = {360--391},
- year = {2005}
+@article{PhelpsTaylor1977,
+ author = {Phelps, Edmund S. and Taylor, John B.},
+ title = {Stabilizing Powers of Monetary Policy under Rational Expectations},
+ journal = {Journal of Political Economy},
+ volume = {85},
+ number = {1},
+ pages = {163--190},
+ year = {1977}
}
@article{SargentWilliams2025,
@@ -4650,82 +4742,43 @@ @article{SargentWilliams2025
note = {Forthcoming}
}
-@article{SargentWilliamsZha2006,
- author = {Sargent, Thomas J. and Williams, Noah and Zha, Tao},
- title = {Shocks and Government Beliefs: The Rise and Fall of American Inflation},
+@article{Orphanides2001,
+ author = {Orphanides, Athanasios},
+ title = {Monetary Policy Rules Based on Real-Time Data},
journal = {American Economic Review},
- volume = {96},
+ volume = {91},
number = {4},
- pages = {1193--1224},
- year = {2006}
-}
-
-@article{Sims1980,
- author = {Sims, Christopher A.},
- title = {Macroeconomics and Reality},
- journal = {Econometrica},
- volume = {48},
- number = {1},
- pages = {1--48},
- year = {1980}
-}
-
-@article{Sims1988,
- author = {Sims, Christopher A.},
- title = {Projecting Policy Effects with Statistical Models},
- journal = {Revista de An\'alisis Econ\'omico},
- volume = {3},
- number = {1},
- pages = {3--20},
- year = {1988}
-}
-
-@article{Sims2001comment,
- author = {Sims, Christopher A.},
- title = {Comment on {Sargent} and {Cogley's} `Evolving Post World War {II} {US} Inflation Dynamics'},
- journal = {NBER Macroeconomics Annual},
- volume = {16},
- pages = {373--379},
+ pages = {964--985},
year = {2001}
}
-@article{SimsZha2006,
- author = {Sims, Christopher A. and Zha, Tao},
- title = {Were There Regime Switches in {U.S.} Monetary Policy?},
- journal = {American Economic Review},
- volume = {96},
- number = {1},
- pages = {54--81},
- year = {2006}
-}
-
-@incollection{Solow1968,
- author = {Solow, Robert M.},
- title = {Recent Controversy on the Theory of Inflation: An Eclectic View},
- booktitle = {Proceedings of a Symposium on Inflation: Its Causes, Consequences, and Control},
- editor = {Rousseaus, Stephen W.},
- publisher = {New York University},
- address = {New York},
- year = {1968}
+@incollection{McLeayTenreyro2019,
+ author = {McLeay, Michael and Tenreyro, Silvana},
+ title = {Optimal Inflation and the Identification of the Phillips Curve},
+ booktitle = {NBER Macroeconomics Annual 2018, Volume 33},
+ editor = {Eichenbaum, Martin and Hurst, Erik and Parker, Jonathan A.},
+ publisher = {University of Chicago Press},
+ address = {Chicago},
+ pages = {199--255},
+ year = {2019}
}
-@article{Stock2001comment,
- author = {Stock, James H.},
- title = {Discussion of {Cogley} and {Sargent} `Evolving Post World War {II} {US} Inflation Dynamics'},
- journal = {NBER Macroeconomics Annual},
- volume = {16},
- pages = {379--387},
- year = {2001}
+@article{CogleySargentConquest2005,
+ author = {Cogley, Timothy and Sargent, Thomas J.},
+ title = {The Conquest of {US} Inflation: Learning and Robustness to Model Uncertainty},
+ journal = {Review of Economic Dynamics},
+ volume = {8},
+ number = {2},
+ pages = {528--563},
+ year = {2005}
}
-@article{StockWatson1998,
- author = {Stock, James H. and Watson, Mark W.},
- title = {Median Unbiased Estimation of Coefficient Variance in a Time-Varying Parameter Model},
- journal = {Journal of the American Statistical Association},
- volume = {93},
- number = {441},
- pages = {349--358},
- year = {1998}
+@book{Bernanke2022,
+ author = {Bernanke, Ben S.},
+ title = {21st Century Monetary Policy: The {Federal} {Reserve} from the Great Inflation to {COVID-19}},
+ publisher = {W. W. Norton and Company},
+ address = {New York},
+ year = {2022}
}
@article{StockWatson2007,
@@ -4738,52 +4791,212 @@ @article{StockWatson2007
year = {2007}
}
-@incollection{Taylor1997comment,
- author = {Taylor, John B.},
- title = {Comment on `America's Peacetime Inflation: The 1970s'},
- booktitle = {Reducing Inflation: Motivation and Strategy},
- editor = {Romer, Christina D. and Romer, David H.},
- publisher = {University of Chicago Press},
- pages = {276--280},
- year = {1997}
+@inproceedings{MurrayAdamsMacKay2010,
+ author = {Murray, Iain and Adams, Ryan P. and MacKay, David J. C.},
+ title = {Elliptical Slice Sampling},
+ booktitle = {Proceedings of the Thirteenth International Conference on
+ Artificial Intelligence and Statistics},
+ series = {Proceedings of Machine Learning Research},
+ volume = {9},
+ pages = {541--548},
+ year = {2010}
}
-@incollection{Tobin1968,
- author = {Tobin, James},
- title = {Discussion},
- booktitle = {Proceedings of a Symposium on Inflation: Its Causes, Consequences, and Control},
- editor = {Rousseaus, Stephen W.},
- publisher = {New York University},
- address = {New York},
- year = {1968}
+@article{KiyotakiWright1989,
+ author = {Kiyotaki, Nobuhiro and Wright, Randall},
+ title = {On Money as a Medium of Exchange},
+ journal = {Journal of Political Economy},
+ volume = {97},
+ number = {4},
+ pages = {927--954},
+ year = {1989},
+ doi = {10.1086/261634}
}
-@article{Whittle1953,
- author = {Whittle, Peter},
- title = {The Analysis of Multiple Stationary Time Series},
- journal = {Journal of the Royal Statistical Society, Series B},
- volume = {15},
- number = {1},
- pages = {125--139},
- year = {1953}
+@article{MarimonMcGrattanSargent1990,
+ author = {Marimon, Ramon and McGrattan, Ellen and Sargent, Thomas J.},
+ title = {Money as a Medium of Exchange in an Economy with Artificially
+ Intelligent Agents},
+ journal = {Journal of Economic Dynamics and Control},
+ volume = {14},
+ number = {2},
+ pages = {329--373},
+ year = {1990},
+ doi = {10.1016/0165-1889(90)90025-C}
}
-@article{Williams2019,
- author = {Williams, Noah},
- title = {Escape Dynamics in Learning Models},
- journal = {Review of Economic Studies},
- volume = {86},
- number = {2},
- pages = {882--912},
- year = {2019}
+@book{Holland1975,
+ author = {Holland, John H.},
+ title = {Adaptation in Natural and Artificial Systems},
+ publisher = {University of Michigan Press},
+ address = {Ann Arbor},
+ year = {1975}
}
-@article{Woodford1990,
- author = {Woodford, Michael},
- title = {Learning to Believe in Sunspots},
- journal = {Econometrica},
- volume = {58},
- number = {2},
- pages = {277--307},
- year = {1990}
+@book{HollandHolyoakNisbettThagard1986,
+ author = {Holland, John H. and Holyoak, Keith J. and Nisbett, Richard E.
+ and Thagard, Paul R.},
+ title = {Induction: Processes of Inference, Learning, and Discovery},
+ publisher = {MIT Press},
+ address = {Cambridge, MA},
+ year = {1986}
+}
+
+@book{Goldberg1989,
+ author = {Goldberg, David E.},
+ title = {Genetic Algorithms in Search, Optimization, and Machine Learning},
+ publisher = {Addison-Wesley},
+ address = {Reading, MA},
+ year = {1989}
+}
+
+@book{Sargent1993,
+ author = {Sargent, Thomas J.},
+ title = {Bounded Rationality in Macroeconomics},
+ publisher = {Oxford University Press},
+ address = {Oxford},
+ series = {The Arne Ryde Memorial Lectures},
+ year = {1993}
+}
+
+@article{KarekenWallace1981,
+ author = {Kareken, John and Wallace, Neil},
+ title = {On the Indeterminacy of Equilibrium Exchange Rates},
+ journal = {Quarterly Journal of Economics},
+ volume = {96},
+ number = {2},
+ pages = {207--222},
+ year = {1981},
+ doi = {10.2307/1882388}
+}
+
+@article{Samuelson1958,
+ author = {Samuelson, Paul A.},
+ title = {An Exact Consumption-Loan Model of Interest with or without the
+ Social Contrivance of Money},
+ journal = {Journal of Political Economy},
+ volume = {66},
+ number = {6},
+ pages = {467--482},
+ year = {1958},
+ doi = {10.1086/258100}
+}
+
+@article{LucasPrescott1971,
+ author = {Lucas, Robert E. and Prescott, Edward C.},
+ title = {Investment Under Uncertainty},
+ journal = {Econometrica},
+ volume = {39},
+ number = {5},
+ pages = {659--681},
+ year = {1971},
+ doi = {10.2307/1909571}
+}
+
+@incollection{MarcetSargent1989hyper,
+ author = {Marcet, Albert and Sargent, Thomas J.},
+ title = {Least Squares Learning and the Dynamics of Hyperinflation},
+ editor = {Barnett, William A. and Geweke, John and Shell, Karl},
+ booktitle = {Economic Complexity: Chaos, Sunspots, Bubbles, and Nonlinearity},
+ publisher = {Cambridge University Press},
+ address = {Cambridge},
+ pages = {119--137},
+ year = {1989}
+}
+
+@article{MarimonSunder1993,
+ author = {Marimon, Ramon and Sunder, Shyam},
+ title = {Indeterminacy of Equilibria in a Hyperinflationary World:
+ Experimental Evidence},
+ journal = {Econometrica},
+ volume = {61},
+ number = {5},
+ pages = {1073--1107},
+ year = {1993},
+ doi = {10.2307/2951494}
+}
+
+@article{BrunoFischer1990,
+ author = {Bruno, Michael and Fischer, Stanley},
+ title = {Seigniorage, Operating Rules, and the High Inflation Trap},
+ journal = {Quarterly Journal of Economics},
+ volume = {105},
+ number = {2},
+ pages = {353--374},
+ year = {1990},
+ doi = {10.2307/2937791}
+}
+
+@article{Brock1974,
+ author = {Brock, William A.},
+ title = {Money and Growth: The Case of Long Run Perfect Foresight},
+ journal = {International Economic Review},
+ volume = {15},
+ number = {3},
+ pages = {750--777},
+ year = {1974},
+ doi = {10.2307/2525739}
+}
+
+@article{Imrohoroglu1993,
+ author = {\.{I}mrohoro\u{g}lu, Selahattin},
+ title = {Testing for Sunspot Equilibria in the {German} Hyperinflation},
+ journal = {Journal of Economic Dynamics and Control},
+ volume = {17},
+ number = {1-2},
+ pages = {289--317},
+ year = {1993},
+ doi = {10.1016/S0165-1889(06)80013-0}
+}
+
+@article{ChenWhite1998,
+ author = {Chen, Xiaohong and White, Halbert},
+ title = {Nonparametric Adaptive Learning with Feedback},
+ journal = {Journal of Economic Theory},
+ volume = {82},
+ number = {1},
+ pages = {190--222},
+ year = {1998},
+ doi = {10.1006/jeth.1998.2426}
+}
+
+@incollection{MarcetMarshall1992,
+ author = {Marcet, Albert and Marshall, David A.},
+ title = {Convergence of Approximate Model Solutions to Rational
+ Expectations Equilibria Using the Method of Parameterized
+ Expectations},
+ booktitle = {Universitat Pompeu Fabra Economics Working Paper},
+ number = {13},
+ year = {1992}
+}
+
+@article{Arifovic1996,
+ author = {Arifovic, Jasmina},
+ title = {The Behavior of the Exchange Rate in the Genetic Algorithm and
+ Experimental Economies},
+ journal = {Journal of Political Economy},
+ volume = {104},
+ number = {3},
+ pages = {510--541},
+ year = {1996},
+ doi = {10.1086/262032}
+}
+
+@book{MinskyPapert1969,
+ author = {Minsky, Marvin and Papert, Seymour},
+ title = {Perceptrons: An Introduction to Computational Geometry},
+ publisher = {MIT Press},
+ address = {Cambridge, MA},
+ year = {1969}
+}
+
+@incollection{Axelrod1987,
+ author = {Axelrod, Robert},
+ title = {The Evolution of Strategies in the Iterated Prisoner's Dilemma},
+ editor = {Davis, Lawrence},
+ booktitle = {Genetic Algorithms and Simulated Annealing},
+ publisher = {Morgan Kaufmann},
+ address = {Los Altos, CA},
+ pages = {32--41},
+ year = {1987}
}
diff --git a/lectures/_toc.yml b/lectures/_toc.yml
index e16fd96..d447729 100644
--- a/lectures/_toc.yml
+++ b/lectures/_toc.yml
@@ -117,11 +117,21 @@ parts:
- file: lq_bewley_complete_markets
- file: lq_robust_smoothing
- file: lq_inventories
+- caption: Bounded Rationality in Macroeconomics
+ numbered: true
+ chapters:
+ - file: bounded_rationality
+ - file: olg_adaptive_money
+ - file: exchange_rate_learning
+ - file: genetic_classifier
+ - file: marimon_mcgrattan_sargent
+ - file: prospects_bounded_rationality
- caption: Phillips Curve Tradeoffs
numbered: true
chapters:
- file: phillips_two_stories
- file: phillips_credibility
+ - file: phillips_credible_policies
- file: phillips_adaptive
- file: phillips_misspecified
- file: phillips_self_confirming
diff --git a/lectures/bounded_rationality.md b/lectures/bounded_rationality.md
new file mode 100644
index 0000000..dfb5582
--- /dev/null
+++ b/lectures/bounded_rationality.md
@@ -0,0 +1,908 @@
+---
+jupytext:
+ text_representation:
+ extension: .md
+ format_name: myst
+ format_version: 0.13
+ jupytext_version: 1.17.1
+kernelspec:
+ display_name: Python 3 (ipykernel)
+ language: python
+ name: python3
+translation:
+ title: 一种奇特的有限理性定义
+ headings:
+ Overview: 概览
+ Overview::The knowledge that rational expectations imputes: 理性预期赋予人们的知识
+ Overview::Sargent's formulation of bounded rationality: 萨金特对有限理性的表述
+ Overview::What the program is good for: 该研究纲领的用处
+ Overview::The rest of the series: 本系列的其余部分
+ Rational expectations as a fixed point: 理性预期作为不动点
+ Rational expectations as a fixed point::A static market: 一个静态市场
+ Rational expectations as a fixed point::Computing the fixed point: 计算不动点
+ Rational expectations as a fixed point::Adaptive expectations: 适应性预期
+ Rational expectations as a fixed point::Muth's inverse problem: 穆思的逆问题
+ Rational expectations as a fixed point::The dynamic analogue: 动态类比
+ Money and prices: 货币与价格
+ Money and prices::Equilibrium: 均衡
+ Money and prices::The bubble: 泡沫
+ Two currencies: 两种货币
+ Where this leaves us: 我们的处境
+ Exercises: 练习
+---
+
+(bounded_rationality)=
+```{raw} jupyter
+
+```
+
+# 一种奇特的有限理性定义
+
+```{index} single: Bounded Rationality
+```
+
+```{contents} Contents
+:depth: 2
+```
+
+## 概览
+
+本讲开启了一个系列,该系列基于 {cite:t}`Sargent1993` 关于宏观经济学中有限理性的著作,反映了20世纪80年代末和90年代初的一种观点。
+
+那些年,苏联解体,原华沙条约国家正在重新安排它们的政治经济体制。
+
+对这些国家来说,那确实是现代动态宏观经济学技术意义上的*体制转变*。
+
+在博弈论、宏观经济学和一般均衡理论方面,两代关于经济动态的研究已经产生了拥抱理性预期假设的理论,这些模型旨在理解人们面对他们已经经历过很多次的反复出现的情形的环境。
+
+20世纪90年代初东欧正在进行的体制转变并非如此。
+
+那里的人们面对的是前所未有的机遇、新的且定义不清的规则,以及每天为弄清楚最终将支配贸易和生产的机制而进行的努力。
+
+拥有良好市场经济模型的经济学家们拥有大量*均衡理论*,描述了一个体系一旦完全适应了一套新的、连贯的规则和预期后将如何运作。
+
+但他们对从苏联体制向市场经济*转变*的动态知之甚少。
+
+他们可能持有关于如何管理这种转变的偏见和轶事,但没有经过实证检验的正式理论。
+
+在这种背景下,一些经济学家冒险进入了 {cite:t}`Sims1980` 所称的非理性预期和有限理性的"荒野"。
+
+其目的一部分是构建转变动态的理论,一部分是理解均衡动态本身的性质,还有一部分是研究那些永远不会稳定下来的体系。
+
+本系列跟随 {cite:t}`Sargent1993` 一起,深入到那片荒野之中。
+
+### 理性预期赋予人们的知识
+
+要看清有限理性*退却*的是什么,先从理性预期说起,它有 **两个** 要求:
+
+1. **个体理性** — 每个人造主体的行为都在其感知约束条件下最大化一个目标函数。
+1. **相互一致性** — 体系中每个人所感知的约束彼此一致。
+
+第二个要求就是穆思的*理性预期*假设。
+
+在一个经济体中,一个人的决策是另一个人约束的一部分,因此一致性要求每个人对其他所有人的决策、决策过程和信念都持有正确的信念。
+
+一致性也正是理性预期力量的来源:如果对感知没有任何限制,一个行为依赖于关于主观信念的任意假设的模型几乎可以产生任何结果。
+
+但请看一看,一旦将模型应用于数据,这一要求赋予人们的东西是什么。
+
+理性预期模型中的主体使用*均衡*概率分布来评估他们的欧拉方程。
+
+而这些正是研究他们的计量经济学家仍在努力估计的分布。
+
+换句话说,这些主体已经以某种方式解决了经济学家才做了一半的推断问题。
+
+### 萨金特对有限理性的表述
+
+萨金特的 **有限理性** 研究纲领保留了个体理性,并以一种特定的方式退却了相互一致性,这一方式受到他对时间序列计量经济学的钟爱所驱动。
+
+萨金特版本的有限理性研究纲领是:
+
+> 我将建立具有"有限理性"主体模型的提议解读为一种呼吁,即通过将理性主体从我们的模型环境中驱逐,代之以行为类似计量经济学家的"人工智能"主体,从而退却理性预期的第二部分(感知的相互一致性)。这些"计量经济学家"进行理论构建、估计和调适,以试图了解在理性预期下他们本已知道的概率分布。
+
+这些主体被塑造得更像那些建立模型的人:他们收集数据,形成理论,进行估计,并进行调适。
+
+```{note}
+在看到萨金特的手稿后,卡内基-梅隆大学的赫伯特·西蒙给萨金特写了一封信,表示他反对萨金特的表述,并建议萨金特不要将他所做的事情称为"有限理性"。
+
+西蒙特别不喜欢萨金特让其模型内的主体表现得像计量经济学家,西蒙认为这种假装是荒谬的。
+```
+
+萨金特的提议使模型建立者的工作变得更难,而不是更容易。
+
+撤回对共同理解环境的假设意味着我们必须用某种东西来取代它,而可供选择的合理选项有很多:
+
+> 这个领域之所以是荒野,是因为研究者一旦决定放弃均衡理论化所提供的约束,就要面对如此多的选择。对均衡理论化的承诺通过要求将人建模为在共同理解的环境中的最优决策者,为他做了许多选择。当我们撤回共同理解环境的假设时,我们必须用某种东西来代替它,而合理的可能性有如此之多。
+
+### 该研究纲领的用处
+
+本书所追求的收益有三种。
+
+有时,一群适应性主体学会*表现得就好像*它们具有理性预期一样,这使得均衡具有了它作为单纯假设时所缺乏的合理性。
+
+有时,适应性主体收敛到众多均衡中的*某一特定*均衡,这就把学习变成了在理性预期均衡中进行**选择**的工具,同时也相应地成为**计算**那些手工难以求解的均衡的一种方法。
+
+而更雄心勃勃的是,适应性动态还带来了一种关于*转变*本身的理论的希望——东欧改革者不得不盲目应对的均衡外调整过程。
+
+正如本系列将会承认的那样,最后这个承诺是实现得最少的;选择和计算方面的收益则更为可靠。
+
+本讲通过最尖锐的情形来处理选择问题:具有**过多均衡**的模型。
+
+当一个理性预期模型有一系列连续的均衡时,经济的物理描述加上均衡概念本身无法确定会发生什么。
+
+必须有其他东西来做出选择,而关于人们如何摸索走向均衡的合理描述正是这个"其他东西"的自然候选者。
+
+我们构建了本系列其余部分会反复回到的两个货币模型示例,两者都具有一系列连续的理性预期均衡:
+
+* 一个关于货币和价格的数量理论模型,其中价格水平只能被确定到一个任意的**泡沫**项;以及
+* 一个双货币版本,其中*汇率完全不受约束*。
+
+在此过程中,我们建立了一些工具——均衡的不动点观点、松弛算法、适应性预期,以及穆思的逆最优预测问题——后续讲座将利用这些工具驱逐理性主体,并代之以适应性主体。
+
+### 本系列的其余部分
+
+* {doc}`olg_adaptive_money` — 萨缪尔森世代交叠货币模型中的适应性家庭。
+ - 最小二乘学习选择了理性预期动态所排斥的低通胀均衡,而实验室中的受试者也走向了同样的方向。
+ - 一个学习菲利普斯曲线的政府作为本讲的结尾。
+* {doc}`exchange_rate_learning` — 本讲的双货币模型,配以适应性主体,其中学习仅通过使汇率依赖于历史来固定汇率,而一个遗传算法经济体反而产生了永不消失的波动。
+* {doc}`genetic_classifier` — 来自霍兰德和联结主义者的候选"大脑"目录:感知器、联想记忆、遗传算法、分类器系统。
+* {doc}`marimon_mcgrattan_sargent` — 分类器系统种群从零开始发现哪种商品将充当货币。
+* {doc}`prospects_bounded_rationality` — 1993年对该研究纲领已取得成就和尚未取得成就的总结,以及关于此后三十年的后记。
+
+让我们从一些导入开始。
+
+```{code-cell} ipython3
+import numpy as np
+import pandas as pd
+import matplotlib.pyplot as plt
+```
+
+## 理性预期作为不动点
+
+### 一个静态市场
+
+考虑一个由大量$n$家相同企业组成的竞争性行业。
+
+每家企业选择产量$x$以最大化
+
+$$
+R(x, p) = p x - c(x),
+$$
+
+其中$c$是一个递增的凸成本函数,价格由一条向下倾斜的逆需求曲线决定
+
+$$
+p = p(nX)
+$$
+
+其中$X$是*平均*企业的产量。
+
+每家企业都是价格接受者和$X$接受者:它将$p$视为给定,并使边际成本等于价格,$p = c'(x)$。
+
+将解写作$x = g(p)$。
+
+代入需求曲线得到**最优反应映射**
+
+```{math}
+:label: br_map
+
+x = g(p(nX)) \equiv h(X),
+```
+
+它将假设的行业平均值$X$映射为一个优化企业相对于该假设值所选择的个体产量。
+
+这是理性预期的第一个组成部分,也是个体理性所能提供的全部内容。
+
+第二个组成部分——一致性——要求每个企业的选择与其假设的平均企业的选择相一致:
+
+```{math}
+:label: static_ree
+
+X = h(X).
+```
+
+一个静态理性预期均衡就是**最优反应映射的一个不动点**。
+
+举一个具体的例子,取一个二次成本函数$c(x) = \tfrac{\gamma}{2} x^2$和一条线性逆需求曲线$p = a - b n X$。
+
+那么$p = c'(x) = \gamma x$给出$x = p / \gamma$,因此
+
+$$
+h(X) = \frac{a - b n X}{\gamma},
+\qquad\text{不动点为}\qquad
+X^* = \frac{a}{\gamma + b n}.
+$$
+
+```{code-cell} ipython3
+a, b, n, γ = 10.0, 0.3, 5, 1.0
+
+def h(X):
+ "Best response of an individual firm to an industry average X."
+ return (a - b * n * X) / γ
+
+X_star = a / (γ + b * n)
+print(f"X* = {X_star}")
+print(f"h(X*) = {h(X_star)}")
+print(f"h'(X) = {-b * n / γ}")
+```
+
+### 计算不动点
+
+现在假设我们想在不求出闭式解的情况下*找到*$X^*$。
+
+一个自然的出发点是**松弛算法**:保留对均衡的估计$X^*_k$,计算对它的最优反应,然后朝该最优反应移动一部分距离,
+
+```{math}
+:label: relaxation
+
+X^*_k = X^*_{k-1} + \lambda\bigl(h(X^*_{k-1}) - X^*_{k-1}\bigr),
+```
+
+其中$\lambda \in (0, 1]$是一个**松弛参数**。
+
+当$\lambda = 1$时,这就是对$h$的简单迭代,即经典的蛛网模型。
+
+```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "Relaxation algorithm paths, undamped and damped"
+ name: fig-br-relaxation
+---
+def relax(λ, X0=1.0, n_iter=12):
+ "Iterate the relaxation algorithm, returning the whole path."
+ path = np.empty(n_iter + 1)
+ path[0] = X0
+ for k in range(n_iter):
+ path[k + 1] = path[k] + λ * (h(path[k]) - path[k])
+ return path
+
+
+fig, axes = plt.subplots(1, 2, figsize=(11, 4))
+
+axes[0].plot(relax(1.0), 'o-', ms=4, lw=1, color='C0',
+ label=r"$\lambda = 1.0$ (cobweb)")
+axes[0].set_title("undamped")
+axes[0].set_ylabel("$X^*_k$")
+
+for i, λ in enumerate((0.8, 0.5, 0.3)):
+ axes[1].plot(relax(λ), 'o-', ms=4, lw=1, color=f'C{i+1}',
+ label=fr"$\lambda = {λ}$")
+axes[1].set_title("damped")
+axes[1].set_ylim(-1, 9)
+
+for ax in axes:
+ ax.axhline(X_star, color='k', lw=0.8, ls='--', label="$X^*$")
+ ax.set_xlabel("$k$")
+ ax.legend(frameon=False, fontsize=9)
+plt.tight_layout()
+plt.show()
+```
+
+朴素的蛛网模型$\lambda = 1$会**发散**,越来越远地振荡偏离它试图寻找的均衡;请注意左侧面板的坐标尺度。
+
+由于$h$是仿射的,迭代 {eq}`relaxation` 的乘子为$1 + \lambda(h' - 1)$,因此它当且仅当以下条件成立时收敛
+
+$$
+\bigl|1 + \lambda(h' - 1)\bigr| < 1
+\qquad\Longleftrightarrow\qquad
+\lambda < \frac{2}{1 + bn/\gamma} .
+$$
+
+```{code-cell} ipython3
+λ_max = 2 / (1 + b * n / γ)
+print(f"converges iff λ < {λ_max}")
+print(f"λ = 1.0 : multiplier = {1 + 1.0 * (-b*n/γ - 1):+.2f} (diverges)")
+print(f"λ = 0.8 : multiplier = {1 + 0.8 * (-b*n/γ - 1):+.2f} (period-2 cycle)")
+print(f"λ = 0.5 : multiplier = {1 + 0.5 * (-b*n/γ - 1):+.2f} (converges)")
+```
+
+在$\lambda = 0.8$时,乘子恰好为$-1$:该方案将$X$映射为$8 - X$,因此永远在$1$和$7$之间弹跳,既不收敛也不发散,这就是右侧面板中未衰减的锯齿形。
+
+充分衰减调整过程能使算法找到均衡;衰减不足则会使算法追逐自己的尾巴。
+
+请记住这一观察结果。
+
+本系列一个反复出现的主题是:*如何*进行调适,决定了你最终会到达哪里,甚至决定了你是否能够到达终点。
+
+### 适应性预期
+
+方程 {eq}`relaxation` 最初是作为一个在迭代次数$k$下运行的算法引入的。
+
+将$k$重新解释为**日历时间**$t$,将$X^*_t$解读为人们*预期*的值,将$X_t = h(X^*_{t-1})$解读为实际发生的情况,它就变成了一种预期形成理论:
+
+```{math}
+:label: adaptive_exp
+
+X^*_t = (1 - \lambda) X^*_{t-1} + \lambda X_t
+ = \lambda \sum_{j=0}^{\infty} (1 - \lambda)^j X_{t-j}.
+```
+
+这就是凯根 {cite:p}`Cagan` 用来研究恶性通货膨胀、以及弗里德曼用来研究消费的**适应性预期**方案。
+
+预期是过去观测值的几何递减分布滞后,用单一自由参数$\lambda$来概括信念。
+
+凯根和弗里德曼把$\lambda$当作一个自由参数,而没有解答为什么会有人以这种方式形成预期这一问题。
+
+### 穆思的逆问题
+
+{cite:t}`Muth1960` 着手通过反过来提出问题来消除$\lambda$这个自由参数。
+
+他没有问某个给定环境暗示什么预测,而是问:**在什么样的环境下,指数平滑会是最优预测?**
+
+他的答案是:当且仅当$X_t$遵循以下过程时,{eq}`adaptive_exp` 才是在每个视界$k$上对$X_{t+k}$的最小二乘预测:
+
+```{math}
+:label: muth_process
+
+X_t = X_{t-1} + \epsilon_t - \theta \epsilon_{t-1},
+```
+
+其中$\{\epsilon_t\}$是一个鞅差序列,并且平滑权重与移动平均系数通过下式相关联:
+
+$$
+\lambda = 1 - \theta .
+$$
+
+```{note}
+萨金特将 {eq}`adaptive_exp` 中的平滑权重和 {eq}`muth_process` 中的移动平均系数都写作$\lambda$。
+
+我们用$\theta$表示后者,以使$\lambda = 1 - \theta$这一关系保持清晰可见。
+
+本讲中另有两个符号身兼两职,都是沿用原书的用法。
+
+在下面的货币模型中,$\lambda$不再是一个增益,而是泡沫项的总增长率$w_1/w_2$,而$\gamma$不再是成本函数的曲率,而是将价格水平与货币供给联系起来的系数。
+
+代码中将它们区分为 `λ_m` 和 `γ_m`。
+```
+
+预测规则从被预测的随机过程中继承了其唯一的参数。
+
+让我们通过数值方法来验证这一点:模拟 {eq}`muth_process`,在一系列权重上运行指数平滑,看看哪个权重预测效果最好。
+
+```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "Forecast MSE of exponential smoothing"
+ name: fig-br-smoothing-mse
+---
+def smoothing_mse(x, λ):
+ "One-step-ahead forecast MSE of exponential smoothing with weight λ."
+ f = np.empty_like(x)
+ f[0] = x[0]
+ for i in range(1, len(x)):
+ f[i] = λ * x[i] + (1 - λ) * f[i - 1]
+ return np.mean((x[1:] - f[:-1]) ** 2)
+
+
+rng = np.random.default_rng(0)
+T = 200_000
+ε = rng.standard_normal(T + 1) # one shock path, shared across all λ
+grid = np.linspace(0.02, 0.98, 97)
+
+fig, ax = plt.subplots(figsize=(7, 4))
+for θ in (0.2, 0.5, 0.8):
+ X = np.cumsum(ε[1:] - θ * ε[:-1]) # Muth's process
+ mse = np.array([smoothing_mse(X, λ) for λ in grid])
+ line, = ax.plot(grid, mse, lw=1.2, label=fr"$\theta = {θ}$")
+ ax.axvline(1 - θ, color=line.get_color(), lw=0.8, ls='--')
+ print(f"θ = {θ}: best λ = {grid[mse.argmin()]:.2f}, 1 - θ = {1 - θ:.2f},"
+ f" min MSE = {mse.min():.4f}")
+ax.set_xlabel(r"smoothing weight $\lambda$")
+ax.set_ylabel("forecast MSE")
+ax.set_ylim(0.9, 3)
+ax.legend(frameon=False)
+plt.show()
+```
+
+每条曲线恰好在其虚线$\lambda = 1 - \theta$处触底,且最小化后的均方误差在模拟噪声范围内等于$\mathbb{V}[\epsilon_t] = 1$——在该权重下的指数平滑*就是*条件期望,没有什么能比它做得更好。
+
+穆思的这一练习是理性预期理念以其后来成为标准形式的首次应用:找出将一种预测方案与它所使用的环境联系起来的限制条件。
+
+它也促使后来的研究者将**预测方案本身**当作定义均衡的对象来处理。
+
+### 动态类比
+
+当我们从静态模型转向动态模型时,正是这种情形发生了。
+
+让每个主体选择一个序列而非单一行动,将有关总量状态的**感知运动规律**视为给定,
+
+$$
+X_t = H(X_{t-1}, u_t),
+$$
+
+其中$\{u_t\}$是独立同分布的。
+
+求解主体的动态规划问题得到个体决策规则
+$x_t = h(x_{t-1}, X_{t-1}, u_t)$。
+
+施加代表性主体确实具有代表性这一条件,即$x_t = X_t$,得出**实际运动规律**
+
+$$
+X_t = h(X_{t-1}, X_{t-1}, u_t) \equiv H^*(X_{t-1}, u_t),
+$$
+
+从而得到从感知运动规律到实际运动规律的一个映射,
+
+```{math}
+:label: t_map
+
+H^* = T(H).
+```
+
+一个动态理性预期均衡是一个不动点$H = T(H)$,与 {eq}`static_ree` 是同一个思想,只是现在这个不动点存在于*函数*空间中,而不是数值空间中。
+
+松弛算法可原样照搬,
+
+```{math}
+:label: t_relaxation
+
+H^*_k = H^*_{k-1} + \lambda\bigl(T(H^*_{k-1}) - H^*_{k-1}\bigr),
+```
+
+只不过现在它是根据预测与实际发生情况之间的差距,对整个预期生成函数进行修订。
+
+本系列中的每一个学习模型都是 {eq}`t_relaxation` 的某个版本,其中增益$\lambda$随时间递减,且修订是由数据驱动的,而不是对$T$的精确求值。
+
+```{seealso}
+{doc}`rational_expectations` 针对一个卢卡斯-普雷斯科特行业模型详细展开了$T$映射,并计算了其不动点。
+
+{doc}`ls_learning` 研究了当主体通过最小二乘法估计感知运动规律,同时其估计又反过来影响它们所处的体系时会发生什么。
+```
+
+## 货币与价格
+
+我们现在构建后续所有内容的两个模型中的第一个。
+
+一个代表性的货币持有者选择从$t$到$t+1$期携带的名义余额$m_t$,以最大化
+
+```{math}
+:label: money_objective
+
+\ln\left(2 w_1 - \frac{m_t}{p_t}\right) + \ln\left(2 w_2 + \frac{m_t}{p^*_{t+1}}\right),
+\qquad w_1 > w_2 > 0,
+```
+
+其中$p_t$是当前价格水平,$p^*_{t+1}$是下一期预期的价格水平。
+
+第一项是今天的消费,由于为获得货币而放弃的实际资源$m_t / p_t$而减少;第二项是明天的消费,由于该货币预期能够购买的商品$m_t / p^*_{t+1}$而增加。
+
+对 {eq}`money_objective` 关于$m_t$求导并整理,得到货币需求
+
+```{math}
+:label: money_demand
+
+\frac{m_t}{p_t} = w_1 - w_2 \frac{p^*_{t+1}}{p_t} .
+```
+
+当货币预期贬值得更快时,实际余额下降。
+
+这是凯根 {cite:p}`Cagan` 用来研究恶性通货膨胀的需求函数的一个版本,它也出现在萨缪尔森 {cite:p}`Samuelson1958` 的世代交叠模型中。
+
+要闭合该模型,我们需要一个关于$p^*_{t+1}$的理论。
+
+假设货币供给以恒定的速率增长,
+
+```{math}
+:label: money_supply
+
+M_{t+1} = \mu M_t ,
+```
+
+并且家庭相信价格水平与货币供给存在如下关系
+
+```{math}
+:label: price_belief
+
+p_t = \gamma M_t + \lambda^t c ,
+```
+
+其中$(\gamma, \lambda, c)$是概括其信念的常数,均为正值,以使价格水平保持为正。
+
+已知$\mu$,家庭预测
+$p^*_{t+1} = \gamma \mu M_t + \lambda^{t+1} c$。
+
+将该预测和 {eq}`price_belief` 代入 {eq}`money_demand`,得到货币需求作为当前货币供给的函数,
+
+```{math}
+:label: money_demand_solved
+
+m_t = \gamma(w_1 - w_2 \mu) M_t + \lambda^t (w_1 - w_2 \lambda) c .
+```
+
+请注意在预期起作用的模型中的一个共同特征:货币的*需求*取决于其*供给*,因为今天的需求取决于明天的预期价格水平,而这被认为取决于明天的货币供给。
+
+### 均衡
+
+令需求等于供给,$m_t = M_t$,将 {eq}`money_demand_solved` 变为一个函数方程,
+
+$$
+M_t = \gamma (w_1 - w_2 \mu) M_t + \lambda^t (w_1 - w_2 \lambda) c ,
+$$
+
+该方程必须在每个日期都成立。
+
+匹配这两项得到
+
+```{math}
+:label: money_ree
+
+\gamma = (w_1 - \mu w_2)^{-1},
+\qquad
+\lambda = \frac{w_1}{w_2},
+\qquad
+c \geq 0 \ \text{ 为任意值} ,
+```
+
+因此均衡价格水平为
+
+```{math}
+:label: money_price_level
+
+p_t = (w_1 - \mu w_2)^{-1} M_t + \left(\frac{w_1}{w_2}\right)^t c .
+```
+
+参数$\gamma$和$\lambda$被确定下来了。
+
+常数$c$却没有被确定。
+
+*每一个$c \geq 0$都是一个理性预期均衡*,并且在每一个均衡中,根据 {eq}`price_belief` 形成的预期总是恰好正确。
+
+让我们数值验证一下,对于我们想尝试的任何$c$,残差确实为零。
+
+```{code-cell} ipython3
+w1, w2, μ, M0 = 2.0, 1.0, 1.5, 1.0
+γ_m, λ_m = 1 / (w1 - μ * w2), w1 / w2
+
+t = np.arange(12)
+M = M0 * μ ** t
+
+def price_path(c):
+ "Equilibrium price level for bubble constant c."
+ return γ_m * M + λ_m ** t * c
+
+for c in (0.0, 0.5, 3.0, 25.0):
+ p = price_path(c)
+ p_star = γ_m * μ * M + λ_m ** (t + 1) * c # forecast of next period's price
+ m = w1 * p - w2 * p_star # money demand
+ print(f"c = {c:5}: max |demand - supply| = {np.max(np.abs(m - M)):.2e}")
+```
+
+在每一个日期,对于每一个$c$,货币需求都恰好等于货币供给。
+
+### 泡沫
+
+由于$w_1 > w_2$,{eq}`money_price_level` 中的第二项以总增长率$\lambda = w_1 / w_2 > 1$增长。
+
+在$c = 0$的均衡中,价格水平与货币供给成比例:这就是教科书形式的数量理论。
+
+在其他每一个均衡中,价格水平都携带一个与货币供给无关、且呈指数增长的成分:一个纯粹投机性的**泡沫**。
+
+```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "Price levels and real balances under bubbles"
+ name: fig-br-bubbles
+---
+fig, axes = plt.subplots(1, 2, figsize=(11, 4))
+
+for c in (0.0, 0.5, 3.0):
+ lab = f"$c = {c}$" + (" (quantity theory)" if c == 0 else "")
+ axes[0].plot(t, price_path(c), 'o-', ms=3, lw=1, label=lab)
+ axes[1].plot(t, M / price_path(c), 'o-', ms=3, lw=1, label=lab)
+
+axes[0].set_yscale('log')
+axes[0].set_ylabel("$p_t$ (log scale)")
+axes[1].set_ylabel("real balances $M_t / p_t$")
+axes[1].axhline(w1 - μ * w2, color='k', lw=0.8, ls='--')
+axes[1].set_ylim(0, 0.6)
+for ax in axes:
+ ax.set_xlabel("$t$")
+ ax.legend(frameon=False, fontsize=9)
+plt.tight_layout()
+plt.show()
+```
+
+左侧面板显示,随着$c$的增大,价格水平与数量理论路径之间的偏离越来越大。
+
+右侧面板显示了这对货币存量的实际价值造成的影响:当$c = 0$时,实际余额恒定在$w_1 - \mu w_2$;而当$c > 0$时,泡沫使实际余额稳步趋向于零。
+
+沿着一条泡沫路径,总通胀率不断攀升,趋向于$w_1 / w_2$,根据 {eq}`money_demand`,这恰恰是实际余额需求消失的速率。
+
+经济体自我去货币化了,纯粹是因为每个人都预期它会这样。
+
+```{code-cell} ipython3
+infl = pd.DataFrame(
+ {f"c = {c}": price_path(c)[1:] / price_path(c)[:-1] for c in (0.0, 0.5, 3.0)},
+ index=pd.Index(t[1:], name="t"),
+).round(3)
+infl
+```
+
+模型中没有任何东西——无论是偏好、技术还是政策——能说明我们究竟处于哪一个$c$之中。
+
+正如这篇论文所说,理性预期"不是一个足以确定结果的、充分限制性的原则"。
+
+## 两种货币
+
+当我们允许存在一种以上货币时,不确定性会变得更严重。
+
+沿用 {cite:t}`KarekenWallace1981`,保持 {eq}`money_demand` 作为货币*总量*的需求,并假设存在两种法定货币,供给量分别为$M_{1t}$和$M_{2t}$,只要它们的回报率相等,它们就是完全替代品:
+
+```{math}
+:label: equal_returns
+
+\frac{p^*_{1,t+1}}{p_{1t}} = \frac{p^*_{2,t+1}}{p_{2t}} .
+```
+
+这种对持有*哪种*货币的无差异性,正是使汇率不确定的原因。
+
+假设人们相信价格水平由下式给出
+
+```{math}
+:label: two_currency_belief
+
+p_{1t} = \gamma_1 M_{1t} + \gamma_2 e M_{2t} + c \lambda^t ,
+\qquad
+p_{2t} = e^{-1} p_{1t} ,
+```
+
+其中$e$是一个恒定的汇率。
+
+要求以货币1单位计价的货币需求等于总供给$M_{1t} + e M_{2t}$,得到
+
+```{math}
+:label: two_currency_ree
+
+\gamma_1 = (w_1 - \mu_1 w_2)^{-1},
+\quad
+\gamma_2 = (w_1 - \mu_2 w_2)^{-1},
+\quad
+\lambda = \frac{w_1}{w_2},
+\quad
+c \geq 0,
+\quad
+e \in [0, \infty) .
+```
+
+这些方程之所以引人注目,正是因为它们所省略的内容。
+
+汇率$e$*完全不受约束*——如果这些方程对某个$e$有解,那么它们对其他任何$e$都有解——而$\gamma_1$和$\gamma_2$的公式根本不涉及$e$。
+
+举最简单的情形:两种货币供给固定,$\mu_1 = \mu_2 = 1$,设$c = 0$。
+
+那么$\gamma_1 = \gamma_2 = (w_1 - w_2)^{-1}$,价格水平为常数。
+
+```{code-cell} ipython3
+w1, w2 = 2.0, 1.0
+H1, H2 = 100.0, 120.0 # fixed supplies of the two currencies
+γ_e = 1 / (w1 - w2)
+
+def two_currency(e):
+ "Price levels and the real allocation at exchange rate e."
+ p1 = γ_e * (H1 + e * H2)
+ p2 = p1 / e
+ supply = H1 + e * H2 # total currency, in units of currency 1
+ demand = (w1 - w2) * p1 # money demand, with p* = p since prices are constant
+ return p1, p2, supply / p1, demand - supply
+
+pd.DataFrame(
+ [two_currency(e) for e in (0.25, 0.5, 1.0, 2.0, 4.0)],
+ index=pd.Index([0.25, 0.5, 1.0, 2.0, 4.0], name="e"),
+ columns=["$p_1$", "$p_2$", "real balances", "excess demand"],
+).round(4)
+```
+
+每一行都是一个均衡。
+
+随着$e$变化,名义价格水平的变动幅度很大。
+
+由于$e$是货币2的一个单位以货币1的单位计价的价值,$e$越大意味着货币2越值钱,这会推高$p_1$并压低$p_2$。
+
+但最后两列讲述了真正的故事:无论$e$为何值,总实际余额都是$w_1 - w_2 = 1$,市场恰好出清。
+
+*在这些均衡中的每一个中,真实配置都是相同的。*模型确定了人们消费什么,以及货币存量能够购买到多少购买力;但它对两种货币之间的兑换比率完全没有任何说法。
+
+这是一个困扰国际货币理论已久的问题的一个尖锐版本,而且它并非一个刀锋边缘的特例:它是一个连续统。
+
+## 我们的处境
+
+我们现在有了两个模型,在这两个模型中,理性预期对我们非常想预测的某件事情保持沉默。
+
+对此有三种应对方式。
+
+第一种是不断给环境增加限制条件,直到均衡变得唯一为止。
+
+第二种是宣称这种不确定性是这个世界的一个真实特征。
+
+第三种——也是本系列所追求的——是问,当我们**用适应性主体取代理性主体**并观察体系走向何方时,会发生什么。
+
+这是一种实质性的改变,而不是技术性的改变。
+
+一个适应性主体并没有被赋予均衡;它拥有信念、修正信念的规则以及初始条件。
+
+那些额外的对象恰恰是理性预期均衡条件未能确定的东西,因此一个由适应性主体构成的体系可以选出一个理性预期无法确定的结果。
+
+本系列的其余部分认真对待这一想法,其结果以一种颇具启发性的方式呈现出好坏参半的局面。
+
+在世代交叠货币经济中,最小二乘学习所选择的均衡,恰恰与理性预期动态所收敛到的均衡相反,而人类实验受试者站在适应性模型这一边。
+
+在上述双货币模型中,适应性主体确实能够固定汇率,但只是通过使其依赖于初始条件来实现的。
+
+学习算法的静止点恰好重现了这种不确定性,而选出某一结果的正是历史那只死去的手。
+
+在一个北原-赖特搜寻经济中,适应性主体学会使用一种交换媒介,并选择*基本面*均衡而非投机性均衡,即便在理论表明投机性均衡是唯一均衡的参数取值下也是如此。
+
+在每一种情形中,算法都提供了均衡概念未能提供的东西。
+
+这究竟是关于经济体的一项发现,还是算法的产物,是本系列不断回到的问题,而 {doc}`prospects_bounded_rationality` 将给出一个定论。
+
+## 练习
+
+```{exercise-start}
+:label: br_ex1
+```
+
+松弛算法 {eq}`relaxation` 对静态市场收敛,当且仅当$\lambda < 2 / (1 + bn/\gamma)$。
+
+比率$bn/\gamma$衡量的是需求斜率相对于成本曲率的大小,因此一个需求陡峭、成本近似线性的市场,正是朴素蛛网模型($\lambda = 1$)表现糟糕的市场。
+
+请通过数值方法验证这一稳定性边界:对于一系列$bn/\gamma$的取值,找出使算法收敛的最大$\lambda$(在一个精细的网格上),并与解析预测进行比较。
+
+```{exercise-end}
+```
+
+```{solution-start} br_ex1
+:class: dropdown
+```
+
+```{code-cell} ipython3
+def converges(slope, λ, n_iter=400, tol=1e-8):
+ """Does the relaxation algorithm converge when h'(X) = -slope?"""
+ a_, X = 10.0, 1.0
+ for _ in range(n_iter):
+ X_new = X + λ * ((a_ - slope * X) - X)
+ if not np.isfinite(X_new) or abs(X_new) > 1e12:
+ return False
+ X, X_prev = X_new, X
+ return abs(X - X_prev) < tol
+
+
+λ_grid = np.linspace(0.01, 1.5, 300)
+rows = []
+for slope in (0.5, 1.0, 1.5, 2.0, 4.0):
+ ok = [λ for λ in λ_grid if converges(slope, λ)]
+ rows.append((slope, max(ok) if ok else np.nan, 2 / (1 + slope)))
+
+pd.DataFrame(rows, columns=["$bn/\\gamma$", "largest $\\lambda$ found",
+ "$2/(1 + bn/\\gamma)$"]).round(3)
+```
+
+数值边界始终位于解析边界的略*下方*,而这个差距并非网格间距造成的;它是该检验方法的一个真实特征。
+
+当$\lambda$接近边界时,乘子趋近于$-1$,因此收敛变得任意缓慢,一个具有固定迭代次数和容差的检验会在真正边界之前就宣告失败。
+
+这个差距的大小是可以预测的:在400次迭代和$10^{-8}$的容差下,只有当$|1 - \lambda(1 + bn/\gamma)|^{400} \lesssim 10^{-8}$时该检验才会通过,即只有当乘子的绝对值低于约$\exp(-18.4/400) = 0.955$时。
+
+```{code-cell} ipython3
+predicted = [(1 + 0.955) / (1 + s_) for s_ in (0.5, 1.0, 1.5, 2.0, 4.0)]
+pd.DataFrame({"$bn/\\gamma$": [0.5, 1.0, 1.5, 2.0, 4.0],
+ "found": [r[1] for r in rows],
+ "predicted by the tolerance": predicted,
+ "true boundary": [r[2] for r in rows]}).round(3)
+```
+
+请注意第一行:当$bn/\gamma < 1$时,边界超过1,因此朴素蛛网模型本身就能收敛,无需任何衰减。
+
+衰减正是在陡峭市场中换取收敛的手段,市场越陡峭,所需的衰减就越多。
+
+```{solution-end}
+```
+
+```{exercise-start}
+:label: br_ex2
+```
+
+均衡 {eq}`money_ree` 的推导过程中并未检验它是否具有经济意义。
+
+请通过找出当货币增长速度快于此时实际余额需求会出现什么问题,来说明货币均衡要求$\mu < w_1 / w_2$。
+
+```{exercise-end}
+```
+
+```{solution-start} br_ex2
+:class: dropdown
+```
+
+沿着$c = 0$的均衡,实际余额恒定为
+$M_t / p_t = \gamma^{-1} = w_1 - \mu w_2$。
+
+只有当$\mu < w_1 / w_2$时,这个值才为正。
+
+如果货币增长得更快,{eq}`money_ree` 仍然会"求解"该函数方程,但它要求家庭持有一个负数量的货币。
+
+```{code-cell} ipython3
+w1, w2 = 2.0, 1.0
+μ_grid = np.array([0.5, 1.0, 1.5, 1.9, 2.0, 2.5])
+
+with np.errstate(divide='ignore'): # γ is infinite exactly at μ = w1/w2
+ table = pd.DataFrame({
+ "$\\mu$": μ_grid,
+ "$\\gamma = (w_1 - \\mu w_2)^{-1}$": 1 / (w1 - μ_grid * w2),
+ "real balances $w_1 - \\mu w_2$": w1 - μ_grid * w2,
+ "monetary equilibrium?": np.where(μ_grid < w1 / w2, "yes", "no"),
+ }).round(3)
+table
+```
+
+在$\mu = w_1 / w_2 = 2$时,对实际余额的需求降为零,$\gamma$则趋于无穷大;超过这一点后,两者都变为负值。
+
+其直觉可以通过 {eq}`money_demand` 来理解:只有当购买力的预期损失$p^*_{t+1} / p_t$小于$w_1 / w_2$时,家庭才会持有货币。
+
+以速率$\mu$增长的货币会产生速率为$\mu$的通货膨胀,因此$\mu \geq w_1 / w_2$会将货币需求推向零,而这恰恰是*泡沫*均衡从下方渐近趋近的同一个边界。
+
+```{solution-end}
+```
+
+```{exercise-start}
+:label: br_ex3
+```
+
+在双货币示例中,我们设定$\mu_1 = \mu_2 = 1$,并发现真实配置在每个汇率下都是相同的。
+
+这一点是特殊的。
+
+请用$\mu_1 \neq \mu_2$——比如$\mu_1 = 1.0$、$\mu_2 = 1.3$,配以$w_1 = 2$、$w_2 = 1$、$M_{1,0} = M_{2,0} = 100$以及$c = 0$——重复这一计算,并在最初几期内计算若干汇率下的总实际余额。
+
+汇率$e$的选择是否仍然不影响真实配置?
+
+```{exercise-end}
+```
+
+```{solution-start} br_ex3
+:class: dropdown
+```
+
+```{code-cell} ipython3
+w1, w2 = 2.0, 1.0
+μ1, μ2 = 1.0, 1.3
+γ1, γ2 = 1 / (w1 - μ1 * w2), 1 / (w1 - μ2 * w2)
+t = np.arange(10)
+M1, M2 = 100.0 * μ1 ** t, 100.0 * μ2 ** t
+
+def real_balances(e):
+ p1 = γ1 * M1 + γ2 * e * M2
+ return (M1 + e * M2) / p1
+
+pd.DataFrame({f"e = {e}": real_balances(e) for e in (0.25, 1.0, 4.0)},
+ index=pd.Index(t, name="t")).round(4)
+```
+
+```{code-cell} ipython3
+fig, ax = plt.subplots(figsize=(7, 4))
+for e in (0.25, 1.0, 4.0):
+ ax.plot(t, real_balances(e), 'o-', ms=3, lw=1, label=f"$e = {e}$")
+ax.axhline(1 / γ1, color='k', lw=0.8, ls='--', label=r"$1/\gamma_1$")
+ax.axhline(1 / γ2, color='gray', lw=0.8, ls=':', label=r"$1/\gamma_2$")
+ax.set_xlabel("$t$")
+ax.set_ylabel("total real balances")
+ax.legend(frameon=False)
+plt.show()
+```
+
+不会。在货币增长率不相等的情况下,汇率会影响真实配置,而且这种配置甚至不再随时间保持恒定。
+
+原因在于$\gamma_1 \neq \gamma_2$:这两种货币被赋予了不同的价值,因为人们预期它们会以不同的速率被稀释,因此货币存量的*构成*就变得重要起来,而$e$正是决定这种构成的因素。
+
+由于货币2增长得更快,无论我们选择哪个$e$,它最终都会主导货币存量,实际余额都会从$e$的起始点开始收敛到$1/\gamma_2$。
+
+因此,固定供给情形下纯粹的名义不确定性是一种刀锋边缘现象,但$e$本身的不确定性却并非如此:表中的每一个$e$仍然都是一个均衡。
+
+```{solution-end}
+```
\ No newline at end of file
diff --git a/lectures/exchange_rate_learning.md b/lectures/exchange_rate_learning.md
new file mode 100644
index 0000000..f8e2e2d
--- /dev/null
+++ b/lectures/exchange_rate_learning.md
@@ -0,0 +1,741 @@
+---
+jupytext:
+ text_representation:
+ extension: .md
+ format_name: myst
+ format_version: 0.13
+ jupytext_version: 1.17.1
+kernelspec:
+ display_name: Python 3 (ipykernel)
+ language: python
+ name: python3
+translation:
+ title: 汇率不确定性、学习与实验
+ headings:
+ Overview: 概览
+ The two-currency economy: 双货币经济
+ The two-currency economy::The indeterminacy, recalled: 回顾不确定性
+ Newton–Raphson learning: 牛顿-拉夫森学习
+ Newton–Raphson learning::The indeterminacy shows up as a singular Hessian: 不确定性表现为一个奇异海森矩阵
+ Newton–Raphson learning::Convergence to a history-dependent exchange rate: 收敛到一个依赖历史的汇率
+ Newton–Raphson learning::The dead hand of history: 历史的死手
+ Newton–Raphson learning::The ghost of indeterminacy: 不确定性的幽灵
+ Evidence from the laboratory: 来自实验室的证据
+ A genetic algorithm economy: 遗传算法经济体
+ A genetic algorithm economy::Volatility that never dies: 永不消退的波动
+ A genetic algorithm economy::The shape of the volatility: 波动的形态
+ Concluding remarks: 结语
+ Exercises: 练习
+---
+
+(exchange_rate_learning)=
+```{raw} jupyter
+
+```
+
+# 汇率不确定性、学习与实验
+
+```{index} single: Bounded Rationality; Exchange Rate Indeterminacy
+```
+
+```{contents} Contents
+:depth: 2
+```
+
+## 概览
+
+在 {doc}`bounded_rationality` 中,我们遇到过一个模型,其中理性预期完全无法确定汇率。
+
+两种法定货币,只要回报率相等就是完美替代品,这使得汇率 $e$ *完全不受限制*:如果均衡条件对某个 $e$ 有解,那么对其他任何 $e$ 也都有解,并且所有情形下的实际配置都是相同的。
+
+本讲座遵循 {cite:t}`Sargent1993` 的思路,将理性主体从这个经济中驱逐出去,代之以适应性主体。
+
+以下内容围绕两个问题展开。
+
+**第一**,学习能否确定汇率?
+
+一个适应性主体拥有信念、修正信念的规则以及初始条件,而这些初始条件恰恰是均衡条件未能确定的东西。
+
+因此,一个由适应性主体组成的系统*可以*在理性预期无法做到的地方选定一个汇率。
+
+我们将看到它确实做到了,但方式非常特殊。
+
+牛顿-拉夫森学习者会收敛到一个**确定的**汇率,而这个汇率完全取决于它们的出发点。
+
+"历史的死手"完成了基本面拒绝完成的锁定任务。
+
+**第二**,不确定性的幽灵是否依然存在?
+
+确实存在。
+
+学习算法的静止点恰好复现了这种不确定性:使算法停止的条件,正是当初使汇率变得自由的那个套利条件。
+
+萨金特将这种使汇率依赖于历史的机制称为确定汇率所依靠的"一根脆弱的芦苇"。
+
+接下来我们转向证据。
+
+{cite:t}`Arifovic1996` 将这一经济体作为付费人类受试者参与的实验室实验来运行。
+
+他们的汇率从未稳定下来。
+
+而当她用**遗传算法**——一个由选择、交叉和变异培育出来的二进制字符串主体群体——替换牛顿-拉夫森学习者时,她得到了一个汇率持续波动的经济体,其频谱与真实的浮动汇率相似。
+
+这种对比——一种收敛的学习规则与一种产生持续波动的学习规则——正是本讲座的落脚点,也是通往 {doc}`marimon_mcgrattan_sargent` 的引桥,在那里遗传算法和分类器系统将完全接管。
+
+让我们先导入一些包。
+
+```{code-cell} ipython3
+import numpy as np
+import pandas as pd
+import matplotlib.pyplot as plt
+from typing import NamedTuple
+```
+
+## 双货币经济
+
+我们采用 {cite:t}`KarekenWallace1981` 不确定性的世代交叠模型化身。
+
+在每个时点 $t$,有 $N$ 个生存两期的主体出生,年轻时禀赋为 $w_1$,年老时禀赋为 $w_2$,且 $w_1 > w_2$。
+
+存在两种供给量固定的法定货币 $H_1$ 和 $H_2$。
+
+一个年轻主体做出两项决策:储蓄多少 $s_t$,以及将这部分储蓄中多大比例 $\lambda_t$ 持有为货币 1(其余部分持有货币 2)。
+
+选择 $(s_t, \lambda_t)$ 的主体实现的终身效用为
+
+```{math}
+:label: xr_utility
+
+U(s_t, \lambda_t)
+= u(w_1 - s_t)
++ u\!\left(w_2 + \lambda_t s_t \frac{p_{1t}}{p_{1,t+1}}
+ + (1 - \lambda_t) s_t \frac{p_{2t}}{p_{2,t+1}}\right),
+```
+
+其中 $p_{it}$ 是以货币 $i$ 计价的价格水平,$p_{it}/p_{i,t+1}$ 是持有货币 $i$ 的总回报率。
+
+我们全程采用 $u(c) = \ln c$。
+
+将每种货币的供给等同于对它的需求,得到价格水平
+
+```{math}
+:label: xr_prices
+
+p_{1t} = \frac{H_1}{\sum_i \lambda_{it} s_{it}},
+\qquad
+p_{2t} = \frac{H_2}{\sum_i (1 - \lambda_{it}) s_{it}},
+```
+
+汇率为 $e_t = p_{1t}/p_{2t}$。
+
+### 回顾不确定性
+
+考虑一个价格恒定的平稳均衡。
+
+那么每种货币的回报为 $p_{it}/p_{i,t+1} = 1$,所以两种回报相等,且 {eq}`xr_utility` 中的组合回报无论 $\lambda$ 为何值都是 $1$。
+
+储蓄决策于是求解 $\max_s \ln(w_1 - s) + \ln(w_2 + s)$,得到
+
+$$
+s^\star = \frac{w_1 - w_2}{2},
+$$
+
+但组合份额 $\lambda$ 是*完全不确定的*:当两种回报相等时,效用与 $\lambda$ 无关。
+
+而 $\lambda$ 恰恰决定了汇率。
+
+由 {eq}`xr_prices`,在公共的 $(s, \lambda)$ 下,
+
+$$
+e = \frac{p_{1}}{p_{2}} = \frac{H_1}{H_2}\cdot\frac{1 - \lambda}{\lambda}.
+$$
+
+每一个 $\lambda \in (0, 1)$ 都是均衡,因此每一个 $e \in (0, \infty)$ 都是均衡。
+
+这正是 {doc}`bounded_rationality` 中的不确定性,现在通过一个未确定的组合选择表现了出来。
+
+本讲座将运行两个参数不同的经济体——萨金特的经济体和阿里福维奇的经济体——因此我们将参数放入一个容器中,并显式传递,而不是让它们作为全局变量随意存在。
+
+```{code-cell} ipython3
+class Params(NamedTuple):
+ w1: float # 年轻时的禀赋
+ w2: float # 年老时的禀赋
+ H1: float # 货币1的固定供给
+ H2: float # 货币2的固定供给
+
+ @property
+ def s_star(self):
+ "理性预期储蓄率,当两种货币回报均为1时。"
+ return (self.w1 - self.w2) / 2
+
+
+kw = Params(w1=20.0, w2=15.0, H1=100.0, H2=120.0)
+
+print(f"rational expectations saving rate s* = {kw.s_star}")
+print(f"any portfolio share λ is an equilibrium, and e = (H1/H2)(1-λ)/λ is free")
+```
+
+## 牛顿-拉夫森学习
+
+现在驱逐理性主体。
+
+遵循 {cite:t}`Sargent1993`,我们将总体分成两类——称为"偶数类"和"奇数类"——因为一个世代的终身效用只有等到该世代年老之后才能被评估。
+
+每一类都携带自己的规则 $(s, \lambda)$,仅根据同一类之前主体的经验来更新。
+
+一个主体根据实现的效用,通过**牛顿-拉夫森**步骤来修正 $(s, \lambda)$:朝着根据 {eq}`xr_utility` 的二阶展开、依据主体实际经历的回报本应能提高效用的方向移动。
+
+用 $g$ 表示梯度,$R$ 表示(负定)海森矩阵的滚动估计,递推式为
+
+```{math}
+:label: xr_newton
+
+\begin{aligned}
+R_{\tau+1} &= R_\tau + \gamma_\tau (H_\tau - R_\tau), \\
+\begin{bmatrix} s \\ \lambda \end{bmatrix}_{\tau+1}
+&= \begin{bmatrix} s \\ \lambda \end{bmatrix}_\tau
+ - \gamma_\tau R_{\tau+1}^{-1}\, g_\tau,
+\end{aligned}
+```
+
+其中 $H_\tau$ 是实现的海森矩阵,$\gamma_\tau$ 是增益。
+
+下面是 {eq}`xr_utility` 关于回报 $R_1 = p_{1t}/p_{1,t+1}$ 和 $R_2 = p_{2t}/p_{2,t+1}$ 的梯度与海森矩阵。
+
+```{code-cell} ipython3
+def grad_hess(p, s, lam, R1, R2):
+ "实现效用 U(s, λ) 的梯度与海森矩阵。"
+ A = lam*R1 + (1 - lam)*R2 # 组合总回报
+ c2 = p.w2 + s*A # 老年消费
+ dR = R1 - R2
+ g = np.array([-1/(p.w1 - s) + A/c2,
+ s*dR/c2])
+ H_ss = -1/(p.w1 - s)**2 - A**2/c2**2
+ H_ll = -s**2 * dR**2 / c2**2
+ H_sl = dR/c2 - s*A*dR/c2**2
+ return g, np.array([[H_ss, H_sl], [H_sl, H_ll]])
+```
+
+### 不确定性表现为一个奇异海森矩阵
+
+看看海森矩阵中 $\lambda$ 那一块。
+
+当两种回报相等,$R_1 = R_2$ 时,所有 $\lambda$-偏导数都消失了:梯度分量 $s(R_1 - R_2)/c_2$ 为零,曲率 $-s^2(R_1 - R_2)^2/c_2^2$ 也是零。
+
+在回报相等时,效用**沿 $\lambda$ 方向是平坦的**。
+
+这正是局部所见的不确定性:没有任何力量将 $\lambda$ 推向任何方向,因此一次牛顿步骤——将梯度除以曲率——在 $\lambda$ 方向上是 $0/0$。
+
+为使算法能够真正移动,我们必须保持 $R$ 可逆。
+
+我们通过在 $\lambda$ 方向引入一个小的**先验曲率** $\kappa$ 来做到这一点。
+
+从经济学角度看,这恰恰是 {cite:t}`Sargent1993` 所援引的迟滞性:正是它作为"历史的死手"使一个原本自由的汇率能够安定下来。
+
+```{code-cell} ipython3
+def newton_learning(p, s0, lam0, T=400, gain=0.3, κ=0.5):
+ """
+ 带有牛顿-拉夫森学习(s, λ)的双货币世代交叠模型,每一类(偶数/奇数)一条规则。
+ κ是不确定λ方向上的先验曲率,用于保持二阶矩矩阵R可逆。
+ """
+ s = np.array(s0, float)
+ lam = np.array(lam0, float)
+ R = [np.array([[-1.0, 0.0], [0.0, -κ]]) for _ in range(2)] # 先验曲率
+ p1 = np.empty(T)
+ p2 = np.empty(T)
+ e = np.empty(T)
+ s_hist = np.empty((T, 2))
+ lam_hist = np.empty((T, 2))
+
+ for t in range(T):
+ j = t % 2 # 在t时年轻的一类
+ p1[t] = p.H1 / (lam[j]*s[j])
+ p2[t] = p.H2 / ((1 - lam[j])*s[j])
+ e[t] = p1[t] / p2[t]
+ if t >= 1: # 在t-1时年轻的一类现在年老了
+ jp = 1 - j
+ R1, R2 = p1[t-1]/p1[t], p2[t-1]/p2[t]
+ g, H = grad_hess(p, s[jp], lam[jp], R1, R2)
+ H[1, 1] -= κ # 保留先验的λ曲率
+ R[jp] = R[jp] + gain*(H - R[jp])
+ step = gain * np.linalg.solve(R[jp], g)
+ s[jp] = np.clip(s[jp] - step[0], 0.1, p.w1 - 0.1)
+ lam[jp] = np.clip(lam[jp] - step[1], 0.02, 0.98)
+ s_hist[t] = s
+ lam_hist[t] = lam
+
+ return e, s_hist, lam_hist
+```
+
+### 收敛到一个依赖历史的汇率
+
+运行两个初始条件不同、其余条件完全相同的实验。
+
+```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "Newton-Raphson learning from two initial conditions"
+ name: fig-xr-learning
+---
+e1, s1, l1 = newton_learning(kw, [3.5, 2.0], [0.35, 0.40])
+e2, s2, l2 = newton_learning(kw, [2.0, 3.5], [0.62, 0.58])
+
+fig, axes = plt.subplots(1, 3, figsize=(14, 4))
+
+axes[0].plot(np.log(e1), 'C0', lw=1.3, label="experiment 1")
+axes[0].plot(np.log(e2), 'C1', lw=1.3, label="experiment 2")
+axes[0].set_title("log exchange rate")
+axes[0].set_xlabel("$t$")
+axes[0].legend(frameon=False)
+
+for s_h, c in [(s1, 'C0'), (s2, 'C1')]:
+ axes[1].plot(s_h[:, 0], color=c, lw=1.0)
+ axes[1].plot(s_h[:, 1], color=c, lw=1.0, ls=':')
+axes[1].axhline(kw.s_star, color='k', lw=0.8, ls='--')
+axes[1].set_title("saving (solid = even, dotted = odd)")
+axes[1].set_xlabel("$t$")
+
+for l_h, c in [(l1, 'C0'), (l2, 'C1')]:
+ axes[2].plot(l_h[:, 0], color=c, lw=1.0)
+ axes[2].plot(l_h[:, 1], color=c, lw=1.0, ls=':')
+axes[2].set_title(r"portfolio share $\lambda$")
+axes[2].set_xlabel("$t$")
+plt.tight_layout()
+plt.show()
+```
+
+两个经济体都收敛了。
+
+储蓄从任一方向都会攀升或下降至理性预期比率 $s^\star = 2.5$,且两类主体的组合份额都汇合到一个共同的值。
+
+但这两个实验收敛到了*不同的汇率*。
+
+它们之间唯一的差别在于组合的起始点。
+
+```{code-cell} ipython3
+for name, e, l in [("experiment 1", e1, l1), ("experiment 2", e2, l2)]:
+ print(f"{name}: e → {e[-1]:.4f}, s → {s1[-1][0]:.3f}, λ → {l[-1][0]:.3f}")
+```
+
+储蓄率是由基本面锁定的;汇率则是由历史锁定的。
+
+### 历史的死手
+
+为了看清历史对结果的支配有多彻底,我们扫描初始组合份额,记录每个经济体最终稳定在哪个汇率上。
+
+```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "Limiting exchange rate against initial portfolio share"
+ name: fig-xr-limits
+---
+λ0_grid = np.linspace(0.15, 0.85, 15)
+e_limits = [newton_learning(kw, [4.0, 4.0], [λ0, λ0])[0][-1] for λ0 in λ0_grid]
+
+fig, ax = plt.subplots(figsize=(7, 4.5))
+ax.plot(λ0_grid, e_limits, 'C0o-', ms=5, label="limiting $e$ from learning")
+ax.plot(λ0_grid, (kw.H1/kw.H2)*(1 - λ0_grid)/λ0_grid, 'k--', lw=1,
+ label=r"$(H_1/H_2)(1-\lambda)/\lambda$")
+ax.set_xlabel(r"initial portfolio share $\lambda_0$")
+ax.set_ylabel("limiting exchange rate $e$")
+ax.legend(frameon=False)
+plt.show()
+```
+
+极限汇率描绘出了*整个理性预期连续统*。
+
+曲线上的每一点都是有效的理性预期均衡;学习纯粹依据初始条件在其中作出选择。
+
+正是在这个意义上,学习使汇率变得确定。
+
+它并没有添加基本面所缺失的某个基本因素。
+
+它把*水平上的不确定性*转化为*对历史的依赖*:经济体最终到达的汇率,就是初始组合信念所蕴含的那个值,被冻结在原地。
+
+### 不确定性的幽灵
+
+为什么汇率会被冻结,而不是移动到某个特定的值?
+
+因为算法在梯度消失的任何地方都会静止下来,而组合梯度 $s(R_1 - R_2)/c_2$ 恰好在 $R_1 = R_2$ 时消失,而这正是使汇率在理性预期下变得自由的那个**套利条件**。
+
+```{code-cell} ipython3
+# at the limit the classes have merged, so prices are constant and R1 = R2 = 1
+s_lim, lam_lim = s1[-1], l1[-1]
+p1_lim = kw.H1 / (lam_lim*s_lim)
+p2_lim = kw.H2 / ((1 - lam_lim)*s_lim)
+print(f"classes converged to a common rule: s = {s_lim.round(4)}, λ = {lam_lim.round(4)}")
+print(f"→ prices constant across periods, so R1 = R2 = 1 (arbitrage holds at the rest point)")
+print(f"→ the λ-gradient is zero for *any* λ, so learning cannot move the exchange rate off "
+ f"wherever history left it")
+```
+
+学习动态的静止点*就是*理性预期均衡本身,汇率也不例外。
+
+不确定性并没有被消除;它被转移到了初始条件之中。
+
+萨金特对这种机制能够承受多大分量说得很直白:
+
+> 换句话说,一个使汇率依赖于历史的机制,似乎是一种形式欠佳的机制。
+
+一个仅由初始信念的偶然性所决定的汇率,是一根脆弱的芦苇。
+
+如果经济体稍有不同——如果学习规则始终在不断试探而不是安定下来——整个构造都可能崩溃。
+
+而这正是接下来两个例证所展示的内容。
+
+## 来自实验室的证据
+
+{cite:t}`Arifovic1996` 用 $w_1 = 11$、$w_2 = 1$、$H_1 = H_2 = 10$ 将这个双货币经济作为一项付费人类受试者实验来实施。
+
+每个年轻的受试者选择一个储蓄率,以及在两种货币之间分配的比例;实验者根据这些选择清算两个货币市场,正如 {eq}`xr_prices` 所描述的那样。
+
+结果与牛顿-拉夫森模拟完全不同。
+
+汇率持续波动,大致在 $0.5$ 到 $2$ 之间的区间内,没有任何安定下来的迹象。
+
+如果说有什么变化的话,波动幅度在各场次实验中反而增大了。
+
+始终收敛到一个常数值的简单牛顿-拉夫森模型,在这方面表现很差。
+
+因此,阿里福维奇为同一个经济体建立了一个不同的模型,其中主体是由遗传算法培育出来的一个**群体**。
+
+## 遗传算法经济体
+
+现在每一类都是 $N = 30$ 个主体组成的群体,每个主体是长度为 $30$ 的**二进制字符串**:前 $20$ 位编码其储蓄率,后 $10$ 位编码其组合份额。
+
+每一期,年轻群体的字符串被解码为 $(s_i, \lambda_i)$ 对,两个货币市场根据 {eq}`xr_prices` 依总量出清,然后读出汇率。
+
+一代之后,当该世代年老时,每个字符串实现的效用 {eq}`xr_utility` 就是其**适应度**,遗传算法据此培育下一代。
+
+```{code-cell} ipython3
+N, L_s, L_lam = 30, 20, 10 # 群体规模;储蓄和组合份额的位数
+
+arifovic = Params(w1=11.0, w2=1.0, H1=10.0, H2=10.0)
+
+def decode(p, pop):
+ "二进制字符串 → (s, λ),其中s在(0, w1)中,λ在(0, 1)中。"
+ ints_s = pop[:, :L_s] @ (1 << np.arange(L_s)[::-1])
+ ints_l = pop[:, L_s:] @ (1 << np.arange(L_lam)[::-1])
+ s = 0.05 + (p.w1 - 0.10) * ints_s / (2**L_s - 1)
+ lam = 0.02 + 0.96 * ints_l / (2**L_lam - 1)
+ return s, lam
+
+def market_prices(p, pop):
+ s, lam = decode(p, pop)
+ return p.H1 / np.sum(lam*s), p.H2 / np.sum((1 - lam)*s)
+
+def fitness(p, pop, R1, R2):
+ "给定每个字符串所经历的回报,计算其实现的终身效用。"
+ s, lam = decode(p, pop)
+ c2 = p.w2 + s*(lam*R1 + (1 - lam)*R2)
+ return np.log(p.w1 - s) + np.log(np.maximum(c2, 1e-9))
+```
+
+这个遗传算法拥有经典的三种算子——按适应度比例进行父代**选择**、单点**交叉**和位翻转**变异**——再加上阿里福维奇引入的一种算子,即**选举算子**:只有当一个子代在最近一期的回报下会比其父代更优时,它才被纳入下一代,否则父代得以存续。
+
+```{code-cell} ipython3
+def genetic_step(p, pop, R1, R2, rng, p_mut=0.033, election=True):
+ "从当前一代培育出新的一代。"
+ fit = fitness(p, pop, R1, R2)
+ weight = fit - fit.min() + 1e-6 # 平移为正值以用于轮盘赌选择
+ new = np.empty_like(pop)
+ for k in range(0, N, 2):
+ i, j = np.searchsorted(np.cumsum(weight), rng.random(2) * weight.sum())
+ parents = np.array([pop[i], pop[j]])
+ cut = rng.integers(1, L_s + L_lam) # 单点交叉
+ kids = parents.copy()
+ kids[0, cut:], kids[1, cut:] = parents[1, cut:], parents[0, cut:]
+ for c in kids: # 变异
+ c[rng.random(L_s + L_lam) < p_mut] ^= 1
+ if election: # 阿里福维奇的选举算子
+ f_kids = fitness(p, kids, R1, R2)
+ f_par = fitness(p, parents, R1, R2)
+ for m in range(2):
+ new[k + m] = kids[m] if f_kids[m] > f_par[m] else parents[m]
+ else:
+ new[k], new[k + 1] = kids
+ return new
+
+
+def genetic_economy(p, T=3000, seed=0, election=True):
+ rng = np.random.default_rng(seed)
+ pops = [rng.integers(0, 2, (N, L_s + L_lam)) for _ in range(2)] # 偶数、奇数
+ p1 = np.empty(T)
+ p2 = np.empty(T)
+ e = np.empty(T)
+ for t in range(T):
+ j = t % 2
+ p1[t], p2[t] = market_prices(p, pops[j])
+ e[t] = p1[t] / p2[t]
+ if t >= 1:
+ jp = 1 - j
+ R1, R2 = p1[t-1]/p1[t], p2[t-1]/p2[t]
+ pops[jp] = genetic_step(p, pops[jp], R1, R2, rng, election=election)
+ return e
+```
+
+### 永不消退的波动
+
+```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "Exchange rate in the genetic-algorithm economy"
+ name: fig-xr-genetic
+---
+e_ga = genetic_economy(arifovic, T=4000, seed=1)
+log_e = np.log(e_ga[500:]) # 舍弃一段预烧期
+
+fig, ax = plt.subplots(figsize=(9, 4))
+ax.plot(log_e, lw=0.5)
+ax.set_xlabel("$t$")
+ax.set_ylabel("$\\log e_t$")
+plt.show()
+```
+
+汇率不断波动,而且持续波动下去。
+
+与牛顿-拉夫森学习者不同,这个遗传群体从未安定下来:变异不断注入新的字符串,市场也不断对其重新定价。
+
+这种波动是一种永久性特征,而不是一种暂态现象。
+
+```{code-cell} ipython3
+early = np.log(e_ga[500:1500]).std()
+late = np.log(e_ga[3000:]).std()
+print(f"standard deviation of log e, early window: {early:.3f}")
+print(f"standard deviation of log e, late window: {late:.3f}")
+print("→ the volatility does not damp out over time")
+```
+
+### 波动的形态
+
+阿里福维奇报告称,她的遗传经济体中的汇率行为几乎就像随机游走——但带有**均值回复**性质,表现为其一阶差分频谱在零频率处出现一个凹陷。
+
+```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "Spectrum and autocorrelation of the exchange rate"
+ name: fig-xr-spectrum
+---
+def averaged_spectrum(x, n_seg=16):
+ "巴特利特平均周期图,用于获得可读的频谱估计。"
+ x = x - x.mean()
+ seg = len(x) // n_seg
+ acc = np.zeros(seg // 2 + 1)
+ for k in range(n_seg):
+ acc += np.abs(np.fft.rfft(x[k*seg:(k+1)*seg]))**2 / seg
+ return np.fft.rfftfreq(seg), acc / n_seg
+
+
+d_log_e = np.diff(np.log(e_ga[500:]))
+d_log_e = np.clip(d_log_e, *np.percentile(d_log_e, [1, 99])) # 缩尾处理极端值
+freq, spec = averaged_spectrum(d_log_e)
+
+def acf(x, K):
+ x = x - x.mean()
+ return np.array([np.sum(x[k:]*x[:len(x)-k]) / np.sum(x*x) for k in range(K)])
+
+
+fig, axes = plt.subplots(1, 2, figsize=(12, 4))
+axes[0].plot(freq[1:], spec[1:], lw=1.2)
+axes[0].axhline(spec[1:].mean(), color='k', ls=':', lw=0.8)
+axes[0].set_title(r"spectrum of $\Delta \log e$")
+axes[0].set_xlabel("frequency")
+
+axes[1].bar(range(25), acf(np.log(e_ga[500:]), 25))
+axes[1].set_title(r"autocorrelation of $\log e$")
+axes[1].set_xlabel("lag")
+plt.tight_layout()
+plt.show()
+```
+
+一阶差分的频谱在零频率附近较低,并向更高频率上升,这是一个*水平*值接近随机游走的序列的特征,但又具有足够的均值回复性,将低频功率压低下去。
+
+水平值的自相关衰减缓慢,如同近单位根序列那样,但它*确实*会衰减,而纯粹的随机游走则不会如此。
+
+```{code-cell} ipython3
+band = slice(1, len(spec)//2)
+print(f"spectral power of Δlog e near zero frequency: {spec[1]:.4f}")
+print(f"average spectral power over low-to-mid band: {spec[band].mean():.4f}")
+print(f"→ pronounced dip at zero frequency (mean reversion): {spec[1] < spec[band].mean()}")
+```
+
+该书的评估是:真实的硬通货货币对浮动汇率的对数差分频谱看起来与此非常相似,只是*没有*零频率处的凹陷;实际汇率更加接近纯粹的随机游走。
+
+阿里福维奇的遗传经济体在完全没有基本面冲击的情况下,仅凭一个不断适应的主体群体对两种本质上完全相同的货币重新定价,就在很大程度上产生了这种效果。
+
+## 结语
+
+对同一个不确定经济体的两个模型给出了两种截然不同的结论。
+
+牛顿-拉夫森学习者会**收敛**——而且正是在收敛的过程中,暴露了这种不确定性,而不是消除了它。
+
+它的静止点是一个使汇率保持自由的套利条件,因此极限汇率就是历史上初始组合信念所蕴含的那个值。
+
+依靠死手来确定,萨金特认为这是一根脆弱的芦苇。
+
+遗传算法经济体则**不会收敛**。
+
+一整个由不断变异刷新、由市场不断重新定价的适应性字符串所组成的群体,产生了永不消退的汇率波动,这种波动模仿了真实浮动汇率的低频行为——而这个经济体根本没有任何基本面扰动。
+
+它们之间的差距并不在于经济学本身,因为经济学是相同的,而在于*适应机制*本身。
+
+一种能够稳定单一主体学习的装置——即只保留改进后代的选举算子——在一个自我指涉的市场内部,结果却维持了波动而不是抑制了波动(见 {ref}`xr_ex2`)。
+
+经济学家采用哪种学习模型,再一次成为有限理性研究纲领被迫摆到台面上来的诸多选择之一。
+
+这套机制——霍兰德的遗传算法及其更为丰富的近亲,分类器系统——正是 {doc}`marimon_mcgrattan_sargent` 的主题,在那里,一个由适应性主体组成的群体不仅要调整组合,还必须*从零开始发现*哪种商品将充当货币。
+
+## 练习
+
+```{exercise-start}
+:label: xr_ex1
+```
+
+牛顿-拉夫森经济体会收敛到一个静止点,在该点上两类主体共享一条公共规则,且套利条件 $R_1 = R_2$ 成立。
+
+通过让两类主体以**非对称**方式起步,验证极限汇率依赖于*整个*初始条件,而不仅仅是初始 $\lambda$ 这一说法。
+
+从若干个初始条件出发运行学习者,其中偶数类和奇数类以不同的组合份额起步,并确认:(a) 两类主体的份额会汇合到一个共同的值;(b) 储蓄会收敛到 $s^\star$;(c) 极限汇率依赖于起始点。
+
+共同的极限 $\lambda$ 是否等于两个初始份额的平均值?
+
+```{exercise-end}
+```
+
+```{solution-start} xr_ex1
+:class: dropdown
+```
+
+```{code-cell} ipython3
+rows = []
+starts = [([3.0, 2.2], [0.30, 0.50]),
+ ([2.2, 3.0], [0.50, 0.30]),
+ ([3.5, 2.0], [0.35, 0.55]),
+ ([2.5, 2.5], [0.40, 0.60])]
+for s0, lam0 in starts:
+ e, s_h, l_h = newton_learning(kw, s0, lam0, T=600)
+ rows.append([str(lam0), l_h[-1][0], l_h[-1][1], s_h[-1].mean(),
+ e[-1], np.mean(lam0)])
+
+pd.DataFrame(rows, columns=["initial λ", "λ even (limit)", "λ odd (limit)",
+ "s (limit)", "e (limit)", "mean of initial λ"]).round(4)
+```
+
+两类主体的份额始终会汇合(第2列和第3列一致),储蓄始终收敛到 $s^\star = 2.5$,且极限汇率在不同起始点之间有所不同——因此它确实依赖于历史。
+
+但共同的极限 $\lambda$ **并不**只是两个初始份额的简单平均:将汇合后的值与最后一列进行比较即可看出。
+
+暂态路径很重要,因为在两类主体仍然存在差异时,回报也存在差异,而这些暂态回报差在套利条件关闭这种运动之前,会推动 $\lambda$ 四处移动。
+
+汇率记录的是整个调整过程的历史,而不仅仅是其起始平均值。
+
+```{solution-end}
+```
+
+```{exercise-start}
+:label: xr_ex2
+```
+
+阿里福维奇的**选举算子**只在一个子代在最新回报下击败其父代时才将其纳入下一代。
+
+在单主体优化中,这样一个过滤器只可能带来好处——它永远不会让一个更差的规则取代一个更好的规则。
+
+但这个经济体是*自我指涉的*:用于衡量适应度的回报本身正是由群体的选择所决定的。
+
+请研究选举算子对汇率波动性的影响。
+
+在若干个种子值下分别运行带有和不带有该算子的遗传经济体,并报告对数汇率的波动性。
+
+该算子将波动性推向哪个方向?你能解释原因吗?
+
+```{exercise-end}
+```
+
+```{solution-start} xr_ex2
+:class: dropdown
+```
+
+```{code-cell} ipython3
+rows = []
+for election in (True, False):
+ sds = [np.log(genetic_economy(arifovic, T=2500, seed=sd,
+ election=election)[500:]).std()
+ for sd in range(6)]
+ rows.append([election, np.mean(sds), np.min(sds), np.max(sds)])
+
+pd.DataFrame(rows, columns=["election operator", "mean s.d. of log e",
+ "min", "max"]).round(3)
+```
+
+选举算子会**提高**波动性——这与它在单主体搜索中的稳定作用恰恰相反。
+
+其机制在于自我指涉性。
+
+在开启该过滤器的情况下,群体只保留在*上一期回报下*击败其父代的后代,因此它会追逐最近有利可图的组合。
+
+但"最近有利可图"取决于汇率,而汇率恰恰会被群体自身的追逐所推动——于是整个群体都涌向最近的赢家,出现过度反应,汇率随之剧烈变动,使另一种组合变得有利可图,于是群体又转而追逐那一种组合。
+
+如果关闭该过滤器,变异会使群体保持分散,各种不同的组合在总量层面 {eq}`xr_prices` 上会部分相互抵消,从而抑制了波动。
+
+一种能够毫不含糊地改善孤立学习者的装置,却可能使一个由学习者组成的市场变得不稳定。
+
+这是一个简明的警示:不能将单主体的直觉简单套用到多主体系统之中——这一主题在 {doc}`marimon_mcgrattan_sargent` 中还会再次出现。
+
+```{solution-end}
+```
+
+```{exercise-start}
+:label: xr_ex3
+```
+
+本讲声称遗传经济体的汇率"接近随机游走,但带有均值回复"。
+
+请将这一说法量化。
+
+将 $\log e_t$ 视为数据,估计一阶自回归模型 $\log e_t = \mu + \phi \log e_{t-1} + \varepsilon_t$ 中的系数,并将其与 $1$(纯随机游走)进行比较。
+
+对若干个种子值分别进行此项估计。
+
+$\phi$ 是否接近但小于 $1$,与一个近单位根、具有均值回复性质的序列相符?
+
+```{exercise-end}
+```
+
+```{solution-start} xr_ex3
+:class: dropdown
+```
+
+```{code-cell} ipython3
+def ar1_coefficient(x):
+ "OLS估计x_t = μ + φ x_{t-1} + ε中的φ。"
+ x0, x1 = x[:-1], x[1:]
+ X = np.column_stack([np.ones_like(x0), x0])
+ μ, φ = np.linalg.lstsq(X, x1, rcond=None)[0]
+ return φ
+
+rows = []
+for sd in range(6):
+ le = np.log(genetic_economy(arifovic, T=3000, seed=sd)[500:])
+ rows.append([sd, ar1_coefficient(le)])
+
+table = pd.DataFrame(rows, columns=["seed", "AR(1) coefficient φ"])
+print(table.round(4).to_string(index=False))
+print(f"\nmean φ across seeds: {table['AR(1) coefficient φ'].mean():.4f}")
+```
+
+估计出的 $\phi$ 始终接近 $1$,但严格小于 $1$——一个高度持续的序列,但仍会回归其均值,而不是永远漂移下去。
+
+这正是频谱所展示的那种"接近随机游走但带有均值回复"的特征:一个纯粹的随机游走会有精确等于 $1$ 的 $\phi$(且零频率处没有凹陷),而略低于 $1$ 的取值则产生了我们所看到的缓慢自相关衰减和低频凹陷。
+
+遗传经济体在没有任何基本面冲击驱动的情况下,仅凭自身就落入了这个近单位根区域——这种持续性完全是由对两种本质上相同的货币进行重新定价的群体动态所制造出来的。
+
+```{solution-end}
+```
\ No newline at end of file
diff --git a/lectures/genetic_classifier.md b/lectures/genetic_classifier.md
new file mode 100644
index 0000000..b39be2a
--- /dev/null
+++ b/lectures/genetic_classifier.md
@@ -0,0 +1,856 @@
+---
+jupytext:
+ text_representation:
+ extension: .md
+ format_name: myst
+ format_version: 0.13
+ jupytext_version: 1.17.1
+kernelspec:
+ display_name: Python 3 (ipykernel)
+ language: python
+ name: python3
+translation:
+ title: 遗传算法与分类器系统
+ headings:
+ Overview: 概览
+ 'The perceptron: an agent as a discriminant function': 感知机:作为判别函数的主体
+ 'Associative memory: the Hopfield network': 联想记忆:霍普菲尔德网络
+ The genetic algorithm: 遗传算法
+ The genetic algorithm::Axelrod's iterated prisoner's dilemma: 阿克塞尔罗德的重复囚徒困境
+ Classifier systems: 分类器系统
+ Classifier systems::A two-armed bandit: 双臂老虎机
+ Evolutionary programming: 演化编程
+ Concluding remarks: 结束语
+ Exercises: 练习
+---
+
+(genetic_classifier)=
+```{raw} jupyter
+
+```
+
+# 遗传算法与分类器系统
+
+```{index} single: Bounded Rationality; Genetic Algorithms
+```
+
+```{contents} Contents
+:depth: 2
+```
+
+## 概览
+
+到目前为止的讲座给适应性主体赋予了相当计量经济学化的大脑。
+
+在 {doc}`olg_adaptive_money` 中,它们运行递归最小二乘法;在 {doc}`exchange_rate_learning` 中,它们针对实现效用采取了牛顿步骤。
+
+在每种情况下,主体都持有一个*参数化*规则并调整其系数。
+
+本讲座考察一套不同的工具箱,即 {cite:t}`Sargent1993` 从约翰·霍兰德(John Holland)关于人工智能的研究以及神经网络的连接主义文献中汲取的工具箱。
+
+这些装置并不调整固定规则的系数。
+
+它们**从经验中发现规则本身**,从一个庞大的可能性空间中挖掘出来。
+
+萨金特的论述框架是,有限理性研究纲领要求我们让模型中充满"表现得像计量经济学家"的主体,而人工智能文献则是这些主体候选大脑的目录:
+
+> 正是从这些文献所积累的方法储备中,我们将挑选出赋予有限理性主体的"大脑"。
+
+我们将巡览其中四种:
+
+1. **感知机**(perceptron),最简单的神经网络,它其实是一个线性判别函数,与"主体作为计量经济学家"这一解读直接相关;
+1. **霍普菲尔德网络**(Hopfield network),一种联想记忆机制,通过在能量函数上向下滚动,从受损输入中恢复存储的模式;
+1. **遗传算法**(genetic algorithm),霍兰德基于种群的搜索方法,我们将其应用于阿克塞尔罗德的重复囚徒困境博弈,观察其发现合作行为;
+1. **分类器系统**(classifier system),霍兰德提出的"大脑即竞争性经济体"的if-then规则集合,其最简单的实例——一个双臂老虎机——已经暴露出一个微妙的局限性。
+
+我们以**演化编程**(evolutionary programming)作为结尾:即一个适应性主体种群,在缓慢趋向均衡的过程中,可被用来*计算*我们无法直接求解的均衡。
+
+这正是 {doc}`marimon_mcgrattan_sargent` 完整实现的思路,因此本讲座是构建下一讲的机制基础。
+
+让我们从一些导入开始。
+
+```{code-cell} ipython3
+import numpy as np
+import matplotlib.pyplot as plt
+```
+
+## 感知机:作为判别函数的主体
+
+最简单的神经网络是单个**感知机**:$k$ 个输入 $x_i$、权重 $w_i$,以及一个输出
+
+$$
+y = S\!\left(\sum_{i=1}^k w_i x_i\right) = S(w^\top x),
+$$
+
+其中 $S$ 是一个将 $\mathbb{R}$ 映射到 $[0, 1]$ 的"压缩函数":可以是阶跃函数,也可以是S型函数
+$S(z) = 1/(1 + e^{-z})$,或任何累积分布函数。
+
+对于固定的权重,感知机是一个**分类器**:当 $w^\top x > 0$ 时它被激活($y = 1$),否则保持静默,因此边界 $w^\top x = 0$ 是一个分隔两个类别的超平面。
+
+训练感知机——选择 $w$ 以最小化 $\sum_t (y_t - S(w^\top x_t))^2$——是一个非线性最小二乘问题,通过我们熟悉的随机梯度递归求解:
+
+$$
+w_t = w_{t-1} + \gamma_t \nabla S(w_{t-1}, x_t)\,(y_t - S(w_{t-1}^\top x_t)),
+$$
+
+这与之前几讲学习模型中运行的 $1/t$ 型更新方式相同。
+
+以书中的例子为例:根据两个标准化特征区分足球运动员和经济学家。
+
+```{code-cell} ipython3
+rng = np.random.default_rng(0)
+n = 100
+economists = rng.multivariate_normal([-1.0, -0.3], [[0.5, 0.1], [0.1, 0.5]], n)
+players = rng.multivariate_normal([1.2, 0.8], [[0.6, 0.0], [0.0, 0.6]], n)
+X = np.vstack([economists, players])
+y = np.r_[np.zeros(n), np.ones(n)]
+X_aug = np.column_stack([np.ones(len(X)), X]) # prepend an intercept
+
+def sigmoid(z):
+ return 1 / (1 + np.exp(-z))
+
+w = np.zeros(3)
+for _ in range(200): # train by stochastic gradient
+ for t in rng.permutation(len(X)):
+ w += 0.1 * (y[t] - sigmoid(X_aug[t] @ w)) * X_aug[t]
+
+print(f"training accuracy: {np.mean((sigmoid(X_aug @ w) > 0.5) == y):.3f}")
+```
+
+{cite:t}`Sargent1993` 强调的要点是,这对计量经济学家来说并非什么异类事物:感知机的决策边界本质上是一个**线性判别函数**。
+
+将感知机的边界法向量与费舍尔线性判别方向进行比较。
+
+```{code-cell} ipython3
+m0, m1 = X[y == 0].mean(0), X[y == 1].mean(0)
+S_within = np.cov(X[y == 0].T) * n + np.cov(X[y == 1].T) * n
+lda_direction = np.linalg.solve(S_within, m1 - m0)
+lda_direction /= np.linalg.norm(lda_direction)
+perceptron_direction = w[1:] / np.linalg.norm(w[1:])
+
+print(f"perceptron boundary normal : {perceptron_direction.round(3)}")
+print(f"Fisher discriminant direction: {lda_direction.round(3)}")
+print(f"cosine similarity: {abs(perceptron_direction @ lda_direction):.4f}")
+```
+
+```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "Perceptron boundary separating the two groups"
+ name: fig-gc-perceptron
+---
+fig, ax = plt.subplots(figsize=(6.5, 5))
+ax.scatter(*economists.T, s=18, c='C0', label="economists ($y=0$)")
+ax.scatter(*players.T, s=18, c='C1', label="football players ($y=1$)")
+xs = np.linspace(X[:, 0].min(), X[:, 0].max(), 2)
+ax.plot(xs, -(w[0] + w[1]*xs) / w[2], 'k-', lw=1.5, label="perceptron boundary")
+ax.set_xlabel("weight (standardized)")
+ax.set_ylabel("salary (standardized)")
+ax.legend(frameon=False)
+plt.show()
+```
+
+这两个方向几乎完全一致。
+
+明斯基和帕佩特著名的批评 {cite:p}`MinskyPapert1969` 恰恰指出,单个感知机*只能*表示线性判别函数,因此无法分离一条直线无法分开的类别。
+
+该领域的复兴源于人们认识到,**分层**感知机——将其输出馈送到进一步的感知机中——可以逼近任意非线性判别函数,这正是 {doc}`back_prop` 所讨论的主题。
+
+就我们的目的而言,感知机是这个目录中与前几讲已经做过的事情最接近的条目:一个参数化规则,通过梯度下降训练,计量经济学家一眼就能认出来。
+
+其余三种大脑则更为陌生。
+
+## 联想记忆:霍普菲尔德网络
+
+第二种装置存储模式并从受损碎片中回忆它们。
+
+将一个模式表示为长度为 $N$ 的 $\pm 1$ 向量。
+
+我们希望存储 $p$ 个模式 $\sigma^1, \ldots, \sigma^p$,使每个模式都是某个动态系统的不动点,从而输入一个受损版本能迅速收敛到最近的存储模式。
+
+**霍普菲尔德网络**(Hopfield network)通过以下动态实现这一点:
+
+$$
+s(t+1) = \operatorname{sgn}(w\, s(t)),
+$$
+
+以及一个由模式本身构建的权重矩阵。
+
+当模式正交时,赫布规则(Hebb's rule)$w = \tfrac{1}{N}\sigma\sigma^\top$ 就足够了;对于仅仅线性独立(相关的)模式,则使用投影规则 $w = \tfrac{1}{N}\sigma V^{-1}\sigma^\top$,其中 $V = \tfrac{1}{N}\sigma^\top \sigma$。
+
+两者都能使每个存储的模式成为精确的不动点。
+
+我们在 $5\times5$ 像素网格上存储五个字母,由于字母共享许多像素,它们是相关的,因此投影规则才是正确的选择。
+
+```{code-cell} ipython3
+PATTERNS = {
+ 'A': ["01110", "10001", "11111", "10001", "10001"],
+ 'E': ["11111", "10000", "11110", "10000", "11111"],
+ 'I': ["11111", "00100", "00100", "00100", "11111"],
+ 'O': ["01110", "10001", "10001", "10001", "01110"],
+ 'T': ["11111", "00100", "00100", "00100", "00100"],
+}
+
+def to_vector(rows):
+ return np.array([1 if c == '1' else -1 for r in rows for c in r])
+
+letters = list(PATTERNS)
+σ = np.array([to_vector(PATTERNS[c]) for c in letters]) # (p, N)
+
+def projection_rule(patterns):
+ "Weight matrix making each (correlated) pattern an exact fixed point."
+ N = patterns.shape[1]
+ Σ = patterns.T # (N, p)
+ V = Σ.T @ Σ / N
+ return Σ @ np.linalg.inv(V) @ Σ.T / N
+
+def energy(w, s):
+ "Hopfield energy; stored patterns are local minima."
+ return -0.5 * s @ w @ s
+
+def recall(w, s0, max_iter=30):
+ "Iterate sgn(w s) to a fixed point."
+ s = s0.copy()
+ for _ in range(max_iter):
+ s_new = np.sign(w @ s)
+ s_new[s_new == 0] = 1
+ if np.array_equal(s_new, s):
+ break
+ s = s_new
+ return s
+
+w_hop = projection_rule(σ)
+fixed = all(np.array_equal(np.sign(w_hop @ σ[i]), σ[i]) for i in range(len(letters)))
+print(f"all stored letters are fixed points: {fixed}")
+print(f"energy of each stored letter: {[round(energy(w_hop, σ[i]), 1) for i in range(len(letters))]}")
+```
+
+每个存储的字母都处于相同的能量 $-N/2$ 处,且都是不动点。
+
+现在破坏几个像素,让网络回落到最近的记忆。
+
+```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "Hopfield recall of letters from corrupted inputs"
+ name: fig-gc-hopfield
+---
+rng = np.random.default_rng(3)
+show = ['E', 'O', 'T']
+
+fig, axes = plt.subplots(len(show), 3, figsize=(6, 6))
+for row, letter in enumerate(show):
+ i = letters.index(letter)
+ corrupt = σ[i].copy()
+ corrupt[rng.choice(25, 4, replace=False)] *= -1 # flip 4 pixels
+ recovered = recall(w_hop, corrupt)
+ for col, (img, title) in enumerate([(σ[i], "stored"),
+ (corrupt, "corrupted"),
+ (recovered, "recovered")]):
+ axes[row, col].imshow(img.reshape(5, 5), cmap='binary', vmin=-1, vmax=1)
+ axes[row, col].set_xticks([])
+ axes[row, col].set_yticks([])
+ if row == 0:
+ axes[row, col].set_title(title)
+plt.tight_layout()
+plt.show()
+```
+
+受损的字母恢复到了它们的原始状态。
+
+恢复的可靠程度取决于损坏了多少内容以及模式之间的相关程度。
+
+一个受损的字母偶尔会落入*另一个*字母的吸引域,或落入设计者从未存储过的虚假混合状态中。
+
+```{code-cell} ipython3
+for n_flip in (3, 5, 7):
+ correct = 0
+ for i in range(len(letters)):
+ for _ in range(40):
+ corrupt = σ[i].copy()
+ corrupt[rng.choice(25, n_flip, replace=False)] *= -1
+ if np.array_equal(recall(w_hop, corrupt), σ[i]):
+ correct += 1
+ print(f"{n_flip} pixels flipped: recovered exactly {correct}/{len(letters)*40}")
+```
+
+这个机制值得点名,因为它在整个研究方案中反复出现。
+
+回忆是**在能量函数上的下降**,其局部最小值就是存储的模式。
+
+输入一个受损的模式会将系统置于能量表面上靠近正确最小值的位置,而动态过程会将其滚动下降。
+
+这正是**模拟退火**(simulated annealing)背后的几何原理,书中提到这一方法用于逃离*不需要的*局部最小值:加入一定量可控的随机"扰动",一开始较大,随时间递减,从而使系统能够跳出浅层虚假吸引域,最终落入一个深层吸引域。
+
+这是牛顿法的随机对应版本,它在分类器系统的"遗传"实验中再次出现。
+
+## 遗传算法
+
+第三种大脑根本不在光滑表面上下降。
+
+霍兰德的**遗传算法**(genetic algorithm)通过演化一个编码为二进制字符串的候选解**种群**(population),在缺乏牛顿法所需光滑性的"崎岖景观"中进行搜索。
+
+给定一个待最大化的适应度函数 $f$ 以及一个由 $N$ 个长度为 $S$ 的字符串组成的种群,每一代应用四个算子:
+
+1. **评估**:计算每个字符串的适应度 $f(x_i)$。
+1. **繁殖**:以与适应度成正比的概率将字符串复制到下一代("有偏轮盘赌")。
+1. **交叉**:将字符串配对,并在随机切割点交换它们的尾部,形成重组了父代片段的子代。
+1. **变异**:以较小的概率 $p_m$ 独立地翻转每个位。
+
+繁殖将种群集中于已经奏效的方案;交叉和变异注入新的候选方案。
+
+交叉是主力算子:它"保留了遗传结构的长片段"——霍兰德所称的*模式*(schemata)——同时仍进行探索,前提是种群足够多样化以供重组。
+
+关键的是,个体字符串本身不学习:每个只存活一代便消亡。
+
+只有*社会*,即种群的序列,才在学习。
+
+这使得遗传算法作为单个大脑的模型显得别扭,而更适合作为种群或市场的模型——这一点在我们讨论到 {doc}`marimon_mcgrattan_sargent` 时至关重要。
+
+### 阿克塞尔罗德的重复囚徒困境
+
+{cite:t}`Axelrod1987` 使用遗传算法演化**重复囚徒困境**(iterated prisoner's dilemma)博弈的策略,该博弈的循环赛著名地由针锋相对(tit-for-tat)策略获胜。
+
+两个玩家各自选择合作(Cooperate)或背叛(Defect);收益结构奖励对合作者的背叛,但惩罚双方互相背叛。
+
+```{code-cell} ipython3
+# 1 = Cooperate, 0 = Defect; T > R > P > S is the prisoner's dilemma ranking
+T_pay, R_pay, P_pay, S_pay = 5, 3, 1, 0
+
+def payoff(a, b):
+ "My payoff when I play a against opponent's b."
+ if a and b: return R_pay # both cooperate
+ if not a and not b: return P_pay # both defect
+ if not a and b: return T_pay # I defect, they cooperate
+ return S_pay # I cooperate, they defect
+```
+
+按照阿克塞尔罗德的方法,策略是一个**70位字符串**:一个策略基于最后三轮的结果作出条件反应。
+
+每轮都是四种结果之一(我的行动、对手的行动),因此三轮会产生 $4^3 = 64$ 种可能的历史,64位指定了每种历史下的行动。
+
+剩余的6位编码了一个假定的博弈前历史,以启动最初的行动。
+
+```{code-cell} ipython3
+def play(gene, opponent, n_rounds=150):
+ "Play a 70-bit genetic strategy against an opponent; return average payoffs."
+ action, premise = gene[:64], gene[64:]
+ hist = list(premise) # last 3 rounds as [my, opp] pairs
+ my_moves, opp_moves = [], []
+ my_total = opp_total = 0
+ for t in range(n_rounds):
+ idx = 0
+ for bit in hist:
+ idx = (idx << 1) | bit # 6 history bits -> index 0..63
+ a = action[idx] # my move from the lookup table
+ b = opponent(my_moves, opp_moves, t)
+ my_total += payoff(a, b)
+ opp_total += payoff(b, a)
+ my_moves.append(a)
+ opp_moves.append(b)
+ hist = [a, b] + hist[:4] # roll the 3-round window
+ return my_total / n_rounds, opp_total / n_rounds
+```
+
+该策略经过培育以便在固定的对手小组中表现出色。
+
+每个小组成员都是历史*就该成员所见*的函数:其第一个参数是对手所打的历史,第二个参数是自己所打的历史。
+
+```{code-cell} ipython3
+panel_rng = np.random.default_rng(2024)
+
+def all_cooperate(opp_moves, own_moves, t): return 1
+def all_defect(opp_moves, own_moves, t): return 0
+def tit_for_tat(opp_moves, own_moves, t): return 1 if t == 0 else opp_moves[-1]
+def grudger(opp_moves, own_moves, t): return 0 if 0 in opp_moves else 1
+def random_play(opp_moves, own_moves, t): return int(panel_rng.random() < 0.5)
+
+PANEL = {'AllC': all_cooperate, 'AllD': all_defect, 'TFT': tit_for_tat,
+ 'Grudger': grudger, 'Random': random_play}
+
+def fitness(gene):
+ "Average payoff against the whole panel."
+ return np.mean([play(gene, opp)[0] for opp in PANEL.values()])
+```
+
+针锋相对策略复制其对手最后一次的行动,而"记仇者"(grudger)一旦对手曾经背叛过一次,便永远背叛下去。
+
+现在开始演化。
+
+```{code-cell} ipython3
+def evolve(N=60, generations=80, p_mut=0.01, seed=0):
+ rng = np.random.default_rng(seed)
+ pop = rng.integers(0, 2, (N, 70))
+ best_hist, mean_hist = [], []
+ for _ in range(generations):
+ fit = np.array([fitness(pop[i]) for i in range(N)])
+ best_hist.append(fit.max())
+ mean_hist.append(fit.mean())
+ weight = fit - fit.min() + 1e-6 # shift positive for roulette
+ nxt = np.empty_like(pop)
+ for k in range(0, N, 2):
+ i, j = np.searchsorted(np.cumsum(weight), rng.random(2) * weight.sum())
+ cut = rng.integers(1, 70) # single-point crossover
+ c1 = np.concatenate([pop[i, :cut], pop[j, cut:]])
+ c2 = np.concatenate([pop[j, :cut], pop[i, cut:]])
+ for c in (c1, c2):
+ c[rng.random(70) < p_mut] ^= 1 # mutation
+ nxt[k], nxt[k+1] = c1, c2
+ pop = nxt
+ fit = np.array([fitness(pop[i]) for i in range(N)])
+ return pop[fit.argmax()], np.array(best_hist), np.array(mean_hist)
+
+
+panel_rng = np.random.default_rng(0) # reset the panel's randomizer
+champion, best_hist, mean_hist = evolve()
+print(f"generation 0: best fitness {best_hist[0]:.3f}")
+print(f"generation {len(best_hist)-1}: best fitness {best_hist[-1]:.3f}")
+```
+
+```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "Population fitness across generations"
+ name: fig-gc-fitness
+---
+fig, ax = plt.subplots(figsize=(7.5, 4))
+ax.plot(best_hist, label="best in population", lw=2)
+ax.plot(mean_hist, label="population mean", lw=2)
+ax.set_xlabel("generation")
+ax.set_ylabel("panel fitness")
+ax.legend(frameon=False)
+plt.show()
+```
+
+适应度攀升,然后保持稳定。
+
+演化究竟发现了什么样的策略?
+
+```{code-cell} ipython3
+print("evolved champion's average payoff against each opponent:")
+for name, opp in PANEL.items():
+ me, them = play(champion, opp)
+ print(f" vs {name:8s}: me = {me:.2f}, opponent = {them:.2f}")
+```
+
+用阿克塞尔罗德的话说,演化得到的策略是*友善且报复性*的。
+
+它与友善的对手(AllC、TFT、Grudger)达成互相合作,拒绝被AllD(全背叛者)欺骗(沦为接近惩罚收益的相互背叛),并且利用了Random(随机策略)。
+
+没有人告诉它要合作;合作行为之所以出现,是因为在包含合作者的小组中,合作是有利可图的。
+
+而且,它在同一小组面前的表现*甚至比针锋相对策略本身还要好*。
+
+```{code-cell} ipython3
+def play_function(strategy, opponent, n=150):
+ "Average payoff of a stateful strategy function against an opponent."
+ mine, theirs, total = [], [], 0
+ for t in range(n):
+ # each side is passed its opponent's history first, then its own
+ a, b = strategy(theirs, mine, t), opponent(mine, theirs, t)
+ total += payoff(a, b)
+ mine.append(a)
+ theirs.append(b)
+ return total / n
+
+panel_rng = np.random.default_rng(0)
+tft_fitness = np.mean([play_function(tit_for_tat, opp) for opp in PANEL.values()])
+print(f"tit-for-tat's own panel fitness: {tft_fitness:.3f}")
+print(f"evolved champion's panel fitness: {best_hist[-1]:.3f}")
+```
+
+这重现了阿克塞尔罗德的核心发现:遗传算法"产生了一个能够赢得该锦标赛的策略,尤其是能够胜过赢得该锦标赛的'针锋相对'策略"。
+
+它在对付合作者时的表现与针锋相对策略非常相似,但其额外的记忆能力使它能从针锋相对策略放过的可利用对手身上榨取更多收益。
+
+差距并不大,也不应该大。
+
+针锋相对策略在这个对手小组面前是一个强有力的策略;额外的记忆所能带来的,是针对那些针锋相对策略能应对但并非最优应对的对手所获得的适度优势。
+
+## 分类器系统
+
+遗传算法演化的是一个其中没有个体学习的种群。
+
+霍兰德的**分类器系统**(classifier system)将同样的演化机制置于*单个主体内部*,作为单一大脑的模型。
+
+萨金特将其描述为霍兰德关于心智作为**竞争性经济体**(competitive economy)的构想:
+
+> 各条语句相互竞争决策的机会。分类器系统以霍兰德称之为竞争性经济体的方式,将遗传算法的要素与其他方面结合起来代表一个大脑。
+
+分类器系统包括:
+
+* **分类器**(classifiers),即以三元字母表 $\{0, 1, \#\}$ 编码的if-then规则,其条件部分与状态匹配,动作部分规定一个行动,其中 $\#$ 是通配符("我不在乎"),使得通用规则能够与特定规则共存。
+* **解码器**(decoder),给定当前状态,找出哪些分类器的条件是匹配的。
+* **拍卖**(auction),选择一个匹配的分类器来执行行动:可以是最强的一个,或以与强度成比例的概率选出的一个。
+* **计账系统**(accounting system),更新每个分类器的**强度**(strength)——即其决策所赚取的净奖励的移动平均值——并在序贯问题中,将奖励从获得报酬的规则*反向*传递给设立这些规则的规则。
+* **遗传算子**(genetic operators),创造新的分类器,进行泛化(添加 $\#$)和特化(移除 $\#$),从而使系统的词汇本身也在演化。
+
+### 双臂老虎机
+
+由布莱恩·亚瑟(Brian Arthur)和卡尔·西蒙(Carl Simon)提出的最简单的分类器系统,玩的是一个**双臂老虎机**(two-armed bandit)。
+
+臂 $i$ 支付一个均值为 $\mu_i$ 的随机奖励,且 $\mu_1 > \mu_2$,但玩家对任何一个分布都一无所知。
+
+该分类器系统持有两条规则——"拉动臂1"和"拉动臂2"——它们的条件总是被满足。
+
+每条规则的强度是该臂所提供收益的移动平均值,选择拉动哪个臂的概率与强度成正比。
+
+这两条规则一开始拥有**相等**的强度:系统对任何一个臂都一无所知,必须从其所经历的奖励中建立自己的估计。
+
+```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "The classifier bandit converges to probability matching"
+ name: fig-gc-bandit
+---
+def two_armed_bandit(μ, σ=0.5, T=20_000, seed=0):
+ "Strength = running average of an arm's payoff; pull ∝ strength."
+ rng = np.random.default_rng(seed)
+ S = np.ones(2) # equal strengths: no prior knowledge
+ τ = np.array([1, 1]) # pull counters
+ pulls = np.empty(T, int)
+ for t in range(T):
+ w = np.maximum(S, 1e-9) # strengths can dip below zero early on
+ i = 0 if rng.random() < w[0] / w.sum() else 1
+ reward = μ[i] + σ * rng.standard_normal()
+ τ[i] += 1
+ S[i] += (reward - S[i]) / τ[i] # running average
+ pulls[t] = i
+ return pulls
+
+pulls = two_armed_bandit([3.0, 1.0])
+frac_best = np.cumsum(pulls == 0) / np.arange(1, len(pulls) + 1)
+
+fig, ax = plt.subplots(figsize=(7.5, 4))
+ax.plot(frac_best, lw=1)
+ax.axhline(0.75, color='k', ls='--', label="probability match $\\mu_1/(\\mu_1+\\mu_2)$")
+ax.axhline(1.0, color='C3', ls=':', label="optimal (always best arm)")
+ax.set_xlabel("$t$")
+ax.set_ylabel("fraction of pulls on the best arm")
+ax.set_ylim(0.5, 1.05)
+ax.legend(frameon=False)
+plt.show()
+```
+
+拉动较好那个臂的比例会收敛,但**不会收敛到一**。
+
+它收敛到 $\mu_1/(\mu_1 + \mu_2)$,即该臂占总预期奖励的份额。
+
+```{code-cell} ipython3
+for μ in ([1.0, 0.5], [2.0, 1.0], [3.0, 1.0]):
+ frac = np.mean(two_armed_bandit(μ)[-5000:] == 0)
+ print(f"μ = {μ}: fraction on best arm = {frac:.3f}, "
+ f"probability match = {μ[0]/(μ[0]+μ[1]):.3f} (optimal = 1.0)")
+```
+
+亚瑟和西蒙证明了这一点:强度比例分类器进行**概率匹配**(probability-matches)。
+
+它按照各臂预期奖励的比例来拉动,而不是专注于最好的那个,因此它永远都在把奖励留在桌面上。
+
+这不是一个需要掩盖的错误;而是关于计账方式的一个教训。
+
+一个分类器系统的好坏,取决于强度分配和传递的方案。
+
+这里使用的朴素规则得到的是概率匹配;更好的规则可以做得更好。
+
+而在*序贯*问题中——一条规则的收益要经过一连串中间决策,才在很久以后才实现——计账必须做一件更困难的事情:为仅仅*建立起*一个未来有利可图的决策的规则给予奖励。
+
+霍兰德为此设计的装置是**桶链**(bucket brigade):每个采取行动的分类器都将其部分强度支付给在它之前刚刚采取行动的那个分类器,即将系统带入当前规则得以行动的状态的那个分类器。
+
+在一条链末端支付的奖励会向后渗透,一条规则接着一条规则,直到从未直接获得报酬的早期设立规则,也因促成了这一结果而获得了强度。
+
+设计这种反向流动是构建分类器系统的核心技艺,也正是 {doc}`marimon_mcgrattan_sargent` 必须做对的地方,以使主体学会今天为了明天的交易而接受货币。
+
+## 演化编程
+
+我们已经巡览了四种大脑。
+
+最后一个想法关乎如何*运用*它们。
+
+贯穿整个这部分内容的一个反复出现的发现是:适应性主体系统,无论多么迟缓,都倾向于收敛到理性预期均衡。
+
+{doc}`olg_adaptive_money` 展示了最小二乘学习者找到一个均衡;{doc}`exchange_rate_learning` 展示了牛顿学习者稳定在(众多均衡中的)一个均衡上。
+
+**演化编程**(evolutionary programming)将这一倾向转化为一种工具。
+
+如果一个适应性主体种群能够可靠地收敛到某个均衡,我们就可以将该种群作为*计算*该均衡的*方法*来运行,尤其是在模型过于复杂以至于无法手工求解的情况下。
+
+萨金特谨慎地说明了这里正在发生和没有发生的事情:
+
+> 适应性主体在"教导"经济学家,正如任何用于求解非线性方程的数值算法都在"教导"数学家一样。当这些主体能够"教导"我们某些东西时,那是因为我们设计它们如此。
+
+这与本系列中反复出现的对偶性相同:一个学习型经济体是一个去中心化的均衡计算过程,而一个均衡计算过程是一个中心化的学习算法。
+
+遗传算法和分类器系统只是比递归最小二乘更丰富的计算引擎,能够搜索崎岖的景观,发现一个好规则的结构,而不仅仅是调整固定规则的系数。
+
+其代表性应用是 {cite:t}`KiyotakiWright1989` 的货币搜索理论模型,其中交换媒介并非被假定,而必须从主体选择如何交易的过程中**涌现**出来。
+
+均衡是一组交易策略和匹配概率,而对于该模型的丰富版本,很难通过解析方法进行刻画。
+
+{cite:t}`MarimonMcGrattanSargent1990` 将霍兰德分类器系统的种群放入该环境中,观察它们收敛到均衡,然后构建了一个没有已知解析解的五种商品版本,让分类器系统提示该均衡的样貌。
+
+## 结束语
+
+这份目录中的四种大脑在其所假设的既定条件上有所不同。
+
+感知机被赋予了一个函数形式,只被要求提供其系数,这就是为什么计量经济学家一眼就能认出它。
+
+霍普菲尔德网络被赋予了模式本身,只被要求回忆它们。
+
+遗传算法只被赋予了一个适应度函数,必须在一个无法计算梯度的空间中进行搜索。
+
+分类器系统被赋予了一套可以书写规则的词汇,必须发现哪些规则值得保留。
+
+沿着这份清单往下走,我们赋予主体的东西越来越少,而要求它发现的东西越来越多,这正是有限理性研究纲领推动我们前进的方向。
+
+在此过程中出现了两个警示,两者都会在下一讲中再次出现。
+
+遗传算法的个体不学习:只有种群在学习,这使得它更像是一个社会的模型,而非一个心智的模型。
+
+而老虎机的例子表明,一个分类器系统的表现是由其*计账方式*——即强度如何分配、竞价和传递——决定的,而不是仅仅因为拥有分类器这一事实本身。
+
+{doc}`marimon_mcgrattan_sargent` 将这套机制组装成从零开始学习使用货币的主体,而这两个警示都直接关系到这些主体最终能够学到什么。
+
+## 练习
+
+```{exercise-start}
+:label: gc_ex1
+```
+
+霍普菲尔德网络的回忆过程是在能量 $E(s) = -\tfrac{1}{2}s^\top ws$ 上的下降过程。
+
+直接验证这一下降过程。
+
+取一个存储的字母,破坏若干像素,并在回忆动态的每一步记录能量。
+
+确认能量从不增加,并且回忆过程会停止在(或低于)存储模式的能量水平。
+
+然后更严重地破坏模式,并对失败情况进行分类:回忆是落在一个*不同*的存储字母上,还是落在一个从未被存储过的*虚假*状态上?
+
+比较各种情况下的能量,并用它来解释为什么会出现错误。
+
+```{exercise-end}
+```
+
+```{solution-start} gc_ex1
+:class: dropdown
+```
+
+```{code-cell} ipython3
+def recall_with_energy(w, s0, max_iter=30):
+ s = s0.copy()
+ trace = [energy(w, s)]
+ for _ in range(max_iter):
+ s_new = np.sign(w @ s)
+ s_new[s_new == 0] = 1
+ trace.append(energy(w, s_new))
+ if np.array_equal(s_new, s):
+ break
+ s = s_new
+ return s, trace
+
+rng = np.random.default_rng(1)
+i = letters.index('E')
+corrupt = σ[i].copy()
+corrupt[rng.choice(25, 5, replace=False)] *= -1
+final, trace = recall_with_energy(w_hop, corrupt)
+
+print(f"energy along the recall path: {[round(e, 2) for e in trace]}")
+print(f"monotonically non-increasing: {all(np.diff(trace) <= 1e-9)}")
+print(f"recovered the intended letter 'E': {np.array_equal(final, σ[i])}")
+```
+
+能量在每一条回忆路径上都单调下降:动态过程只能向下移动,这就是为什么网络总是停止在一个局部最小值处。
+
+现在更严重地破坏模式,并将结果分为三类:预期的字母、一个*不同的*存储字母,以及一个虽是不动点但从未被教授过的虚假状态。
+
+```{code-cell} ipython3
+def classify_outcome(final, i):
+ if np.array_equal(final, σ[i]):
+ return "intended"
+ if any(np.array_equal(final, σ[k]) for k in range(len(letters))):
+ return "wrong letter"
+ return "spurious"
+
+counts = {"intended": 0, "wrong letter": 0, "spurious": 0}
+wrong_energy, spurious_energy = [], []
+for i in range(len(letters)):
+ for seed in range(200):
+ r = np.random.default_rng(1000*i + seed)
+ corrupt = σ[i].copy()
+ corrupt[r.choice(25, 7, replace=False)] *= -1
+ final = recall(w_hop, corrupt)
+ kind = classify_outcome(final, i)
+ counts[kind] += 1
+ if kind == "wrong letter":
+ wrong_energy.append(energy(w_hop, final))
+ elif kind == "spurious":
+ spurious_energy.append(energy(w_hop, final))
+
+print(f"7-pixel corruptions ({sum(counts.values())} trials): {counts}")
+print(f"stored patterns all have energy {energy(w_hop, σ[0]):.1f}")
+print(f"wrong-letter results: energy in [{min(wrong_energy):.1f}, {max(wrong_energy):.1f}]")
+print(f"spurious results: energy in [{min(spurious_energy):.1f}, {max(spurious_energy):.1f}]")
+```
+
+两种失败模式都会出现,而且在这种破坏程度下,虚假状态在数量上要更为常见一些。
+
+一个**错误字母**的结果恰好处于存储的能量 $-N/2$ 上:这种破坏将起点推过了一个盆地边界,进入了另一个同样深的存储记忆的吸引域。
+
+能量下降会收敛到*某个*最小值,但当相关模式的吸引域相互交错时,无法保证收敛到*最近*的那个。
+
+一个**虚假**的结果是设计者从未打算得到的局部最小值,通常是存储模式的混合,作为存储规则的一种副产品被创造出来。
+
+其能量范围从比存储模式浅一些,一直到与之完全相同的深度,因此仅凭深度本身并不能判断某个记忆是否是我们所要求的那种。
+
+一些最深的虚假状态是符号反转:因为 $\operatorname{sgn}(w(-s)) = -\operatorname{sgn}(ws)$,只要 $\sigma$ 是不动点,向量 $-\sigma$ 就以同样的能量成为不动点,因此无论我们是否希望如此,网络都会存储每个字母的"底片"版本。
+
+由于所有这些都是真正的局部最小值,向下的动态过程无法从任何一个中逃脱。
+
+这两点都是网络不完美的原因,也正是为什么**模拟退火**很重要:加入递减的随机扰动,能让系统在稳定下来之前跳出一个浅层虚假吸引域,或跨越一个盆地边界,而纯粹的能量下降永远做不到这一点。
+
+```{solution-end}
+```
+
+```{exercise-start}
+:label: gc_ex2
+```
+
+遗传算法的探索来自两个算子:交叉和变异。
+
+书中指出,交叉"是该算法的核心",而单独的变异"是注入多样性的一种糟糕机制"。
+
+在阿克塞尔罗德的博弈中检验这一论断。
+
+编写一个仅使用变异的变体(每个子代都是从一个被选中的父代经变异得到的副本,不进行交叉),并将其达到的适应度与完整算法进行比较。
+
+书中谨慎地指出了*何时*单独的变异是弱势的:"当变异率设定为非常低的值时,单独的变异是一种注入多样性的糟糕机制。"
+
+因此在**低**变异率下运行比较,此时交叉必须承担探索的重任。
+
+```{exercise-end}
+```
+
+```{solution-start} gc_ex2
+:class: dropdown
+```
+
+```{code-cell} ipython3
+def evolve_no_crossover(N=60, generations=60, p_mut=0.005, seed=0):
+ "Mutation-only variant: each child is a mutated copy of one selected parent."
+ rng = np.random.default_rng(seed)
+ pop = rng.integers(0, 2, (N, 70))
+ best_hist = []
+ for _ in range(generations):
+ fit = np.array([fitness(pop[i]) for i in range(N)])
+ best_hist.append(fit.max()) # recorded exactly as in `evolve`
+ weight = fit - fit.min() + 1e-6
+ parents = np.searchsorted(np.cumsum(weight), rng.random(N) * weight.sum())
+ nxt = pop[parents].copy()
+ nxt[rng.random((N, 70)) < p_mut] ^= 1
+ pop = nxt
+ return best_hist[-1]
+
+rows = []
+for seed in range(5):
+ panel_rng = np.random.default_rng(seed)
+ with_x = evolve(N=60, generations=60, p_mut=0.005, seed=seed)[1][-1]
+ panel_rng = np.random.default_rng(seed)
+ without_x = evolve_no_crossover(seed=seed)
+ rows.append((seed, with_x, without_x))
+
+print(f"{'seed':>4} {'with crossover':>16} {'mutation only':>16}")
+for sd, a, b in rows:
+ print(f"{sd:>4} {a:>16.3f} {b:>16.3f}")
+print(f"{'mean':>4} {np.mean([r[1] for r in rows]):>16.3f} "
+ f"{np.mean([r[2] for r in rows]):>16.3f}")
+```
+
+在低变异率下,交叉的优势很明显:在几乎每个随机种子上,它都达到了比仅使用变异的变体更高的适应度。
+
+原因正是书中所给出的。
+
+在变异很少的情况下,仅使用变异的种群只能每次通过一次罕见的位翻转来缓慢前进,并且随着选择过程不断复制其少数最优字符串,多样性很快就会丧失。
+
+而交叉则重组了整段已经在不同字符串中被证明有用的*片段*:将一种应对某个对手的良好方式,与应对另一个对手的良好方式拼接在一起。
+
+它注入了大规模、结构化的变异,同时保留了适应度已经青睐的模式,这正是霍兰德将其置于该算法核心地位的原因。
+
+```{note}
+在较高的变异率下,这一差距会缩小,甚至可能消失:当变异本身就能产生大量多样性时,交叉的贡献就不那么关键了。
+
+尝试用 `p_mut=0.02` 重新运行该比较,看看这一效应如何缩小。
+
+交叉是否具有决定性作用,既取决于变异率,也取决于编码方式是否将有用的构建模块与连续的位段对齐——对于这个基于历史索引的策略表而言,它只做到了部分对齐。
+```
+
+```{solution-end}
+```
+
+```{exercise-start}
+:label: gc_ex3
+```
+
+双臂老虎机分类器进行概率匹配,这是次优的:它会永远以固定的比例继续拉动较差的那个臂。
+
+一个自然的修正方法是,随着系统信心的增强,使拍卖变得更加"贪婪"。
+
+用一个softmax函数替代与强度成比例的选择规则:
+
+$$
+\pi_1 = \frac{e^{\beta S_1}}{e^{\beta S_1} + e^{\beta S_2}},
+$$
+
+其中 $\beta$ 控制贪婪程度($\beta \to \infty$ 时总是选择更强的那个臂)。
+
+实现它,并展示长期而言拉动最佳臂的比例如何依赖于 $\beta$。
+
+$\beta$ 需要多大,才能使分类器从概率匹配转向最优策略?
+
+```{exercise-end}
+```
+
+```{solution-start} gc_ex3
+:class: dropdown
+```
+
+```{code-cell} ipython3
+def bandit_softmax(μ, β, σ=0.5, T=20_000, seed=0):
+ rng = np.random.default_rng(seed)
+ S = np.ones(2) # equal strengths, as before
+ τ = np.array([1, 1])
+ pulls = np.empty(T, int)
+ for t in range(T):
+ p1 = 1 / (1 + np.exp(-β * (S[0] - S[1])))
+ i = 0 if rng.random() < p1 else 1
+ reward = μ[i] + σ * rng.standard_normal()
+ τ[i] += 1
+ S[i] += (reward - S[i]) / τ[i]
+ pulls[t] = i
+ return np.mean(pulls[-5000:] == 0)
+
+μ = [3.0, 1.0]
+print(f"probability match target = {μ[0]/(μ[0]+μ[1]):.3f}, optimal = 1.000\n")
+for β in (0.5, 1.0, 2.0, 4.0, 8.0):
+ frac = bandit_softmax(μ, β)
+ print(f"β = {β:>4}: fraction on best arm = {frac:.3f}")
+```
+
+随着 $\beta$ 的增大,分类器摆脱了概率匹配,转而集中拉动较好的那个臂,逐渐逼近总是拉动该臂的最优策略。
+
+这个练习具体展示了为什么*计账和拍卖规则*——而不仅仅是拥有分类器这一事实——决定了一个分类器系统的表现好坏。
+
+亚瑟和西蒙的强度比例规则只是一个谱系上的一个点;贪婪规则则位于谱系的另一端。
+
+真正的分类器系统,包括 {doc}`marimon_mcgrattan_sargent` 中的那个系统,都是刻意选择其拍卖和强度更新规则的,正是因为这一选择才是区分一个仅仅进行概率匹配的系统与一个学会良好行动的系统的关键所在。
+
+```{solution-end}
+```
\ No newline at end of file
diff --git a/lectures/marimon_mcgrattan_sargent.md b/lectures/marimon_mcgrattan_sargent.md
new file mode 100644
index 0000000..c76ba3c
--- /dev/null
+++ b/lectures/marimon_mcgrattan_sargent.md
@@ -0,0 +1,1843 @@
+---
+jupytext:
+ text_representation:
+ extension: .md
+ format_name: myst
+ format_version: 0.13
+ jupytext_version: 1.17.1
+kernelspec:
+ display_name: Python 3 (ipykernel)
+ language: python
+ name: python3
+translation:
+ title: 人工智能主体中作为交换媒介的货币
+ headings:
+ Overview: 概览
+ The Kiyotaki-Wright environment: Kiyotaki-Wright 环境
+ The Kiyotaki-Wright environment::Two equilibria: 两个均衡
+ Classifier systems: 分类器系统
+ Classifier systems::The auction: 拍卖
+ Classifier systems::The bucket brigade: 消桶传递法
+ Implementation: 实现
+ Implementation::Describing an economy: 描述一个经济体
+ Implementation::The agent: 主体
+ Implementation::The simulation: 模拟
+ Implementation::Reporting: 报告
+ 'Economy A1.1: does a medium of exchange emerge?': 经济体 A1.1:交换媒介会出现吗?
+ 'Economy A2.1: when theory predicts speculation': 经济体 A2.1:当理论预测投机时
+ 'Economy B.1: a different production pattern': 经济体 B.1:一种不同的生产模式
+ The genetic algorithm: 遗传算法
+ 'Economy A1.2: learning from random rules': 经济体 A1.2:从随机规则中学习
+ 'Economy C: fiat money': 经济体 C:法定货币
+ 'Economy D: five goods, five types': 经济体 D:五种商品,五种类型
+ Concluding remarks: 结束语
+ Exercises: 练习
+---
+
+(marimon_mcgrattan_sargent)=
+```{raw} jupyter
+
+```
+
+# 人工智能主体中作为交换媒介的货币
+
+```{index} single: Bounded Rationality; Money as a Medium of Exchange
+```
+
+```{contents} Contents
+:depth: 2
+```
+
+## 概览
+
+Kiyotaki 和 Wright {cite}`KiyotakiWright1989` 研究了一个不存在需求双重巧合的经济。
+
+只有当某种商品被接受不是因为它被需要,而是因为它可以在以后被转手时,交易才能发生。
+
+扮演这种角色的商品就是**交换媒介**。
+
+Kiyotaki 和 Wright 在假设主体完全理性——主体知道其交易伙伴之间商品的分布并对此做出最优反应——的前提下,刻画了这种经济的*平稳纳什均衡*。
+
+Marimon、McGrattan 和 Sargent {cite}`MarimonMcGrattanSargent1990` 提出了一个不同的问题。
+
+假设我们彻底抛弃理性假设,取而代之的是一群**人工智能主体**,他们最初遵循任意的、甚至是随机的经验法则,并且仅仅通过记录过去哪些法则奏效来调整这些法则。
+
+这样的主体会*学会*使用交换媒介吗?
+
+而当理性预期模型存在多个均衡时,如果确实会出现某个均衡,那会是哪一个?
+
+Marimon、McGrattan 和 Sargent 使用的学习装置是 John Holland 的**分类器系统** {cite}`Holland1975,HollandHolyoakNisbettThagard1986`:一个由“如果-那么”规则组成的种群、一个决定哪条规则起作用的拍卖机制,以及一个对导致良好结果的规则记入贷方、对导致不良结果的规则记入借方的记账系统。
+
+作为一个可选项,**遗传算法**会繁育新规则并淘汰旧规则。
+
+本讲重建了他们的计算实验。
+
+我们将复现的主要发现是:
+
+1. 从强度相等的规则的完全枚举开始,甚至从随机生成的规则开始,持有量和交易模式都会收敛到 Kiyotaki-Wright 模型的一个平稳纳什均衡。
+1. 当 Kiyotaki-Wright 模型同时存在*基本*均衡与*投机*均衡时,人工智能主体会选择基本均衡——即储藏成本最低的商品充当货币流通的那个均衡。
+1. 一种本质上毫无价值、储藏无成本的物品——**法定货币**——会被必须自己发现其用处的主体在交易中接受。
+1. 同样的机制在一个有五种商品和五种类型的经济中也能奏效,而作者对此经济的均衡并没有解析刻画:该算法被用作*均衡发现装置*。
+
+让我们从一些导入开始。
+
+```{code-cell} ipython3
+import numpy as np
+import pandas as pd
+import matplotlib.pyplot as plt
+from dataclasses import dataclass
+```
+
+## Kiyotaki-Wright 环境
+
+存在三种类型的主体,用 $i = 1, 2, 3$ 索引,以及三种商品,用 $k = 1, 2, 3$ 索引。
+
+类型 $i$ 的主体只从消费商品 $i$ 中获得效用。
+
+他拥有生产商品 $i^*$ 的技术,其中 $i^* \neq i$。
+
+在 Kiyotaki 和 Wright 的*模型 A* 中,生产模式是“维克塞尔三角”
+
+| 类型 $i$ | 生产 $i^*$ | 消费 |
+|---|---|---|
+| 1 | 2 | 1 |
+| 2 | 3 | 2 |
+| 3 | 1 | 3 |
+
+因此不存在需求的双重巧合:拥有你想要的东西的主体从不想要你拥有的东西。
+
+所有商品都是不可分的,每个主体从一期到下一期恰好能储藏一单位恰好一种商品。
+
+储藏商品 $k$ 一期的成本为 $s_k$,且
+
+$$
+s_3 > s_2 > s_1 > 0 .
+$$
+
+每种类型有相等数量 $A_i$ 的主体,因此总数为 $A = 3 A_i$。
+
+每期每个主体都会被随机配对到恰好一个其他主体,不考虑类型。
+
+记 $x_{at}$ 为主体 $a$ 在 $t$ 期携带的商品,$\rho_t(a)$ 为与 $a$ 配对的主体。
+
+主体 $a$ 的**交易前状态**是一对
+
+$$
+z_{at} = \bigl(x_{at},\; x_{\rho_t(a)t}\bigr) .
+$$
+
+每期每个主体依次做出两个决策。
+
+**第一**,在看到 $z_{at}$ 之后,他决定是否提出交易,
+
+$$
+\lambda_{at} = \begin{cases}
+1 & \text{提出用 } x_{at} \text{ 换取 } x_{\rho_t(a)t} \\
+0 & \text{拒绝}
+\end{cases}
+$$
+
+当且仅当 $\lambda_{at} \lambda_{\rho_t(a)t} = 1$ 时交易才会发生,因此交易后的持有量为
+
+```{math}
+:label: mms_posttrade
+
+x^+_{at} = (1 - \lambda_{at}\lambda_{\rho_t(a)t}) x_{at}
+ + \lambda_{at}\lambda_{\rho_t(a)t} x_{\rho_t(a)t} .
+```
+
+**第二**,他决定是否消费他手中留下的东西,
+
+$$
+\gamma_{at} = \begin{cases}
+1 & \text{消费 } x^+_{at} \\
+0 & \text{将 } x^+_{at} \text{ 带入 } t+1
+\end{cases}
+$$
+
+如果他消费,他会立即生产商品 $f(a) = i^*$,并将其带入 $t+1$。
+
+因此
+
+```{math}
+:label: mms_lom
+
+x_{a,t+1} = \gamma_{at} f(a) + (1 - \gamma_{at}) x^+_{at} .
+```
+
+单期净收益为
+
+```{math}
+:label: mms_payoff
+
+U_a(\gamma_{at}) =
+\gamma_{at}\bigl[u_i(x^+_{at}) - s(f(a))\bigr]
+- (1 - \gamma_{at})\, s(x^+_{at}) ,
+```
+
+其中当 $k = i$ 时 $u_i(k) = u_i > 0$,否则 $u_i(k) = 0$。
+
+请注意,主体*可以*消费他不想要的商品;他只是从中获得零效用,同时仍然要生产并支付储藏 $f(a)$ 的费用。
+
+不消费也不是免费的——它花费 $s(x^+_{at})$。
+
+学习该做哪一种是问题的一部分。
+
+```{note}
+Kiyotaki 和 Wright 按照预期贴现效用对收益流进行排序。
+
+Marimon、McGrattan 和 Sargent 则假设每个主体关心他的**长期平均**效用。
+
+这一点很重要:正如我们将看到的,这正是分类器系统内部的记账系统是由滚动平均值构建的原因。
+```
+
+### 两个均衡
+
+由于主体的收益取决于其他主体的行为,该模型可以存在多个平稳均衡。
+
+Kiyotaki 和 Wright 用一组概率来刻画均衡,其中对我们最有用的是
+
+$$
+\pi^h_{it}(k) = \text{类型 } i \text{ 的主体在 } t \text{ 时持有商品 } k \text{ 的概率} .
+$$
+
+在**基本均衡**中,商品 1——储藏成本最低的商品——成为普遍的交换媒介:
+
+* 类型 1 主体总是持有商品 2(他们自己生产的),并用其交换商品 1;
+* 类型 3 主体总是持有商品 1,并用其交换商品 3;
+* 类型 2 主体一半时间持有商品 1,一半时间持有商品 3。
+
+所以均衡持有概率为
+
+| | $k=1$ | $k=2$ | $k=3$ |
+|---|---|---|---|
+| $i=1$ | 0 | 1 | 0 |
+| $i=2$ | 0.5 | 0 | 0.5 |
+| $i=3$ | 1 | 0 | 0 |
+
+类型 2 主体接受商品 1,即使他们从不消费它:他们把它当作货币使用。
+
+在**投机均衡**中,类型 1 主体还额外接受商品 3——储藏成本*最高*的商品——因为他们预期能很快将其换成商品 1。
+
+Kiyotaki 和 Wright 证明,在贴现率趋于零的极限下,如果
+
+```{math}
+:label: mms_fundcond
+
+s_3 - s_2 > \bigl(\pi^h_1(3) - \pi^h_1(2)\bigr)\tfrac{1}{3} u_1 ,
+```
+
+基本均衡是唯一的平稳均衡;而当不等式反向时,投机均衡是唯一的均衡。
+
+因此,只要将 $u_1$ 提高到足够程度,就能使模型的预测从基本均衡转变为投机均衡。
+
+下面的经济体 A2 恰恰做到了这一点,这是检验适应性主体能否跟踪理性预期预测的良好测试。
+
+## 分类器系统
+
+主体并不被赋予一个策略。
+
+他被赋予的是**一群候选规则**以及一种记分方法。
+
+**分类器**是三元字母表 $\{0, 1, \#\}$ 中的一个字符串,被分为**条件**部分和**动作**部分,其中 $\#$ 表示“无所谓”。
+
+{cite}`Goldberg1989` 是一本关于分类器系统与遗传算法的专著。
+
+商品用二进制编码,条件用三进制编码,因此对于三种商品,两个位置就足够了:
+
+| 编码 | 含义 |
+|---|---|
+| `1 0` | 商品 1 |
+| `0 1` | 商品 2 |
+| `0 0` | 商品 3 |
+| `0 #` | 不是商品 1 |
+| `# 0` | 不是商品 2 |
+| `# #` | 任何商品 |
+
+**交换分类器**是一个长度为 7 的字符串:两个位置表示自己的持有物,两个位置表示伙伴的持有物,最后一个二进制数字表示动作($1$ = 提议交易,$0$ = 拒绝)。
+
+例如
+
+```
+1 0 0 0 1 -> 1 "如果我持有商品1,我的伙伴持有商品3,则提议交易"
+1 0 # # -> 0 "如果我持有商品1,拒绝与任何人交易"
+```
+
+共有 $6 \times 6 \times 2 = 72$ 种不同的交换分类器,这是状态 $z_{at}$ 上可定义的所有规则的*完全枚举*。
+
+**消费分类器**是一个长度为 4 的字符串:两个位置表示交易后的持有量 $x^+_{at}$,一个动作数字($1$ = 消费)。
+
+这样的分类器共有 $6 \times 2 = 12$ 个。
+
+### 拍卖
+
+日期 $t$ 时附加在分类器 $e$ 上的是一个**强度** $S^a_e(t)$。
+
+给定状态 $z_{at}$,令
+
+$$
+M_e(z_{at}) = \{e : z_{at} \text{ 匹配 } e \text{ 的条件部分}\}
+$$
+
+为条件被满足的分类器集合。
+
+起作用的分类器是最强的匹配分类器,
+
+```{math}
+:label: mms_auction
+
+e_t(z_{at}) = \arg\max\,\{S^a_e(t) : e \in M_e(z_{at})\} ,
+```
+
+消费分类器也以同样的方式从 $M_c(z_{at})$ 中选出。
+
+### 消桶传递法
+
+强度通过一个 Holland 称之为*消桶传递法*的内部支付系统进行更新。
+
+只有赢得拍卖的分类器才会改变其强度,因此我们为每个分类器附加一个计数器 $\tau^a_e(t)$,记录它到 $t$ 为止赢得拍卖的次数,初始值为 1。
+
+一个匹配的分类器 $e$ 会投出其强度中的以下比例
+
+$$
+b_1(e) = b_{11} + b_{12}\sigma_e ,
+\qquad
+\sigma_e = \frac{1}{1 + \text{ } e \text{ 中 } \#\text{ 的数量}}
+$$
+
+作为出价,对消费分类器同样有 $b_2(c) = b_{21} + b_{22}\sigma_c$。
+
+由于 $\sigma_e$ 随特殊性增加而升高,在强度相等的情况下,特定规则会压过泛化规则。
+
+支付按以下方式流动。
+
+* 外部收益 $U_a(\gamma_t)$ 支付给 $t$ 时获胜的**消费**分类器。
+* $t$ 时获胜的消费分类器将其出价支付给 $t$ 时获胜的**交换**分类器,因为后者创造了使其得以行动的状态。
+* $t$ 时获胜的交换分类器将其出价支付给 $t-1$ 时获胜的**消费**分类器,因为后者建立了状态 $z_{at}$。
+
+正是这条链将消费所获得的奖励反向传递给使消费成为可能的交易。
+
+由此产生的运动法则为
+
+```{math}
+:label: mms_strengthc
+
+S^a_{c,\tau_c(t)} = S^a_{c,\tau_c(t)-1}
+ - \frac{1}{\tau_c(t)-1}\Bigl[(1 + b_2(c))S^a_{c,\tau_c(t)-1}
+ - \sum_e I^a_e(t) b_1(e) S^a_{e,\tau_e(t)} - U_a(\gamma_{ct})\Bigr]
+```
+
+```{math}
+:label: mms_strengthe
+
+S^a_{e,\tau_e(t)+1} = S^a_{e,\tau_e(t)}
+ - \frac{1}{\tau_e(t)}\Bigl[(1 + b_1(e))S^a_{e,\tau_e(t)}
+ - \sum_c I^a_c(t) b_2(c) S^a_{c,\tau_c(t)}\Bigr]
+```
+
+其中 $I^a_e(t)$ 和 $I^a_c(t)$ 是赢得 $t$ 时拍卖的指示变量。
+
+{eq}`mms_strengthc` 中的时序值得仔细研读,因为它正是跨期传递奖励的方式。
+
+在日期 $t$ 更新的消费分类器是在 $t-1$ 时获胜的那一个:它收集自己决策所赚取的外部收益,还收集来自*现在*获胜的交换分类器的出价 $b_1(e)S_e$,因为正是它创造了后者行动的机会。
+
+而交换分类器则在当期由随后的消费分类器支付。
+
+因此每笔出价都沿着这条链
+
+$$
+\cdots \;\to\; c_{t-1} \;\to\; e_t \;\to\; c_t \;\to\; e_{t+1} \;\to\; \cdots
+$$
+
+向后传递一步,消费时收取的收益就这样一环一环地反向渗透,回到使其成为可能的交易那里。
+
+提出交易但*未*得到回应的交换分类器不会被扣费,其计数器也不会前进:主体从被拒绝的提议中学不到任何东西。
+
+```{note}
+方程 {eq}`mms_strengthc`-{eq}`mms_strengthe` 使强度成为过去净收入的**累积平均值**而非累积总量,而后者正是 Holland 最初规范所使用的。
+
+这正是使强度收敛的创新之处。
+
+由于增益为 $1/\tau$,这些是随机逼近递归式,因此任何极限点都必须满足
+
+$$
+\mathbb{E}\Bigl[(1 + b_2(c))S_c - \sum_e I_e b_1(e) S_e - U(\gamma_c)\Bigr] = 0,
+\qquad
+\mathbb{E}\Bigl[(1 + b_1(e))S_e - \sum_c I_c b_2(c) S_c\Bigr] = 0 .
+$$
+
+Marimon、McGrattan 和 Sargent 将满足这些方程的一组强度定义为*平稳的*,并将在平稳强度下赢得拍卖的规则恰好是支持均衡行为的规则的情形定义为受支持的平稳纳什均衡。
+```
+
+正如论文中一样,给定类型的所有主体共享一个分类器系统;这在节省计算量的同时,使得所有类型 $i$ 的主体同时进行实验。
+
+## 实现
+
+我们将分类器种群存储为一组并行的 NumPy 数组,而不是对象列表,这使我们能够通过单次向量化比较找到所有匹配的规则。
+
+通配符 $\#$ 用 $-1$ 表示。
+
+```{code-cell} ipython3
+WILD = -1
+
+# binary codes for goods, one row per good
+CODES = {
+ 3: np.array([[1, 0], [0, 1], [0, 0]]),
+ 4: np.array([[1, 0], [0, 1], [0, 0], [1, 1]]), # good 4 = fiat money
+ 5: np.array([[1, 0, 0], [0, 1, 0], [0, 0, 1],
+ [1, 1, 0], [1, 0, 1]]),
+}
+
+# the six conditions expressible with two trits (see the table above)
+CONDS_2 = np.array([[1, 0], [0, 1], [0, 0],
+ [0, WILD], [WILD, 0], [WILD, WILD]])
+
+
+def rule_string(cond, action):
+ "Print a classifier the way the paper does."
+ body = ''.join('#' if b == WILD else str(int(b)) for b in cond)
+ return f"{body} -> {int(action)}"
+```
+
+```{code-cell} ipython3
+class Rules:
+ """
+ A population of classifiers held as parallel arrays.
+
+ cond[i] condition part of rule i, entries in {0, 1, WILD}
+ action[i] action part of rule i, in {0, 1}
+ strength[i] S_i, a running average of net receipts
+ used[i] the counter tau_i, initialized at 1
+ traded[i] number of times rule i actually executed a trade
+
+ """
+
+ def __init__(self, cond, action, strength=None):
+ self.cond = np.asarray(cond, dtype=np.int64)
+ self.action = np.asarray(action, dtype=np.int64)
+ n = len(self.action)
+ self.strength = (np.zeros(n) if strength is None
+ else np.asarray(strength, float).copy())
+ self.used = np.ones(n, dtype=np.int64)
+ self.traded = np.zeros(n, dtype=np.int64)
+
+ @property
+ def n(self):
+ return len(self.action)
+
+ @property
+ def length(self):
+ return self.cond.shape[1]
+
+ def matched(self, state):
+ "Boolean mask of rules whose condition is satisfied by state."
+ return np.all((self.cond == WILD) | (self.cond == state), axis=1)
+
+ def specificity(self, i):
+ "sigma_i = 1 / (1 + number of wildcards)."
+ return 1.0 / (1.0 + np.count_nonzero(self.cond[i] == WILD))
+
+ def replace(self, i, cond, action, strength, used=1, traded=0):
+ "Overwrite rule i."
+ self.cond[i] = cond
+ self.action[i] = action
+ self.strength[i] = strength
+ self.used[i] = used
+ self.traded[i] = traded
+```
+
+完全枚举将每个条件与两个动作配对。
+
+随机种群则均匀地抽取条件和动作,这就是不完全枚举经济体的初始方式。
+
+```{code-cell} ipython3
+def complete_rules(pair):
+ "All two-trit rules; pair=True for exchange rules, False for consumption rules."
+ if pair:
+ conds = np.array([np.concatenate([a, b])
+ for a in CONDS_2 for b in CONDS_2])
+ else:
+ conds = CONDS_2.copy()
+ return Rules(np.repeat(conds, 2, axis=0), np.tile([0, 1], len(conds)))
+
+
+def random_rules(n, length, rng):
+ return Rules(rng.integers(-1, 2, size=(n, length)),
+ rng.integers(0, 2, size=n),
+ rng.random(n) * 0.1)
+```
+
+拍卖 {eq}`mms_auction` 挑选最强的匹配规则,平局时按位置打破。
+
+```{code-cell} ipython3
+def auction(rules, state):
+ "Index of the strongest rule matching state, or -1 if none matches."
+ idx = np.flatnonzero(rules.matched(state))
+ if idx.size == 0:
+ return -1, idx
+ return idx[np.argmax(rules.strength[idx])], idx
+```
+
+即使没有遗传算法,也需要两个算子。
+
+**创造**在当前状态完全没有规则匹配时触发:一个冗余或较弱的规则会被一个条件恰好等于刚观察到的状态的规则覆盖,动作随机抽取。
+
+**多样化**在所有匹配规则都要求相同动作时触发:会种入一条动作相反的规则,以便可以尝试并评分这一替代选项。
+
+两者都保持种群规模不变。
+
+```{code-cell} ipython3
+def create(rules, state, rng):
+ "No rule matches state, so overwrite the weakest of the most redundant rules."
+ groups = {}
+ for i in range(rules.n):
+ groups.setdefault(tuple(rules.cond[i]), []).append(i)
+ biggest = max(groups.values(), key=len)
+ group = biggest if len(biggest) > 1 else range(rules.n)
+ j = min(group, key=lambda i: rules.strength[i])
+ rules.replace(j, state, rng.integers(0, 2), rules.strength.mean())
+ return j
+
+
+def diversify_simple(rules, matches):
+ "If every matched rule takes the same action, plant the opposite action."
+ if len(set(rules.action[matches])) > 1:
+ return
+ weak = matches[np.argmin(rules.strength[matches])]
+ rules.replace(weak, rules.cond[weak], 1 - rules.action[matches[0]],
+ rules.strength[matches].mean())
+```
+
+### 描述一个经济体
+
+`Economy` 将物理环境的参数与学习算法的设置一起收集起来。
+
+字段 `method` 从论文使用的三种方案中选择一种:
+
+* `'enumerate'` —— 从强度全为零的完全枚举规则开始,不使用遗传算法;
+* `'ga3'` —— 从随机规则开始,用单点交叉和变异来演化;
+* `'ga4'` —— 从随机规则开始,用一种*泛化*交叉来演化,用于最大的经济体。
+
+如果存在法定货币,它是最后一种商品:储藏成本为零,不带来效用,且不能被消费。
+
+```{code-cell} ipython3
+@dataclass
+class Economy:
+ name: str
+ produces: np.ndarray # produces[i] = good produced by type i
+ storage_costs: np.ndarray # one entry per good
+ u: float = 100.0 # utility from own consumption good
+ n_agents_per_type: int = 50
+ method: str = 'enumerate' # 'enumerate' | 'ga3' | 'ga4'
+ n_trade_rules: int = 72
+ n_consume_rules: int = 12
+ b_trade: tuple = (0.025, 0.025) # (b11, b12)
+ b_consume: tuple = (0.25, 0.25) # (b21, b22)
+ n_fiat: int = 0 # units of fiat money injected at t = 0
+ start: str = 'random' # initial holdings: 'random' | 'production'
+ pcross: float = 0.6
+ pmutation: float = 0.01
+
+ @property
+ def n_types(self):
+ return len(self.produces)
+
+ @property
+ def n_goods(self):
+ return len(self.storage_costs)
+
+ @property
+ def n_agents(self):
+ return self.n_types * self.n_agents_per_type
+
+ @property
+ def fiat(self):
+ return self.n_goods > self.n_types
+
+ @property
+ def n_bits(self):
+ return CODES[self.n_goods].shape[1]
+
+ def code(self, good):
+ return CODES[self.n_goods][good]
+```
+
+### 主体
+
+`Agent` 代表一种*类型*:它拥有一个交换规则种群和一个消费规则种群,并记住上一期哪条消费规则获胜,以便这一期的交换规则能够对它进行支付。
+
+```{code-cell} ipython3
+class Agent:
+
+ def __init__(self, i, econ, rng):
+ self.i, self.econ, self.rng = i, econ, rng
+ if econ.method == 'enumerate':
+ self.trade = complete_rules(pair=True)
+ self.consume = complete_rules(pair=False)
+ else:
+ self.trade = random_rules(econ.n_trade_rules, 2 * econ.n_bits, rng)
+ self.consume = random_rules(econ.n_consume_rules, econ.n_bits, rng)
+ self.pending = None # last period's consumption rule, awaiting settlement
+
+ def decide(self, rules, state, specialize=False):
+ "Run the auction on `state`, applying the operators the method calls for."
+ win, matches = auction(rules, state)
+ if win < 0: # creation
+ j = create(rules, state, self.rng)
+ return rules.action[j], j
+ if self.econ.method == 'enumerate':
+ diversify_simple(rules, matches)
+ elif self.econ.method == 'ga4':
+ diversify_clone(rules, matches, win)
+ if specialize:
+ specialize_winner(rules, win, self.econ.pmutation, self.rng)
+ if self.econ.method != 'ga3':
+ win, matches = auction(rules, state) # the population may have changed
+ return rules.action[win], win
+
+ def trade_decision(self, own, partner, specialize=False):
+ state = np.concatenate([self.econ.code(own), self.econ.code(partner)])
+ return self.decide(self.trade, state, specialize)
+
+ def consume_decision(self, good, specialize=False):
+ return self.decide(self.consume, self.econ.code(good), specialize)
+
+ def update(self, e, c, payoff, active):
+ """
+ The bucket brigade laws of motion for strengths.
+
+ `e` and `c` index the winning exchange and consumption rules and `active`
+ records whether the exchange rule's action was actually carried out.
+
+ """
+ b11, b12 = self.econ.b_trade
+ b21, b22 = self.econ.b_consume
+ T, C = self.trade, self.consume
+ b1 = b11 + b12 * T.specificity(e)
+ b2 = b21 + b22 * C.specificity(c)
+
+ if active: # exchange rule: pays b1, receives b2 * S_c
+ τ = T.used[e]
+ T.used[e] += 1
+ T.traded[e] += int(T.action[e] == 1)
+ T.strength[e] -= ((1 + b1) * T.strength[e] - b2 * C.strength[c]) / τ
+
+ # Now settle last period's consumption rule. Its update waits a period
+ # because only now is the second of its two receipts known: it collects
+ # the payoff its own decision earned *and* the bid of the exchange rule
+ # winning today, whose chance to act it created.
+ if self.pending is not None:
+ p, u_prev = self.pending
+ inflow = b1 * T.strength[e] if active else 0.0
+ b2p = b21 + b22 * C.specificity(p)
+ tau_p = C.used[p]
+ C.used[p] += 1
+ C.strength[p] -= ((1 + b2p) * C.strength[p] - inflow - u_prev) / tau_p
+
+ self.pending = (c, payoff)
+```
+
+### 模拟
+
+每一期,所有 $A$ 个主体被打乱并配成 $A/2$ 对;每对进行交易、消费和强度更新;然后,在不完全枚举经济体中,遗传算子运行。
+
+作者代码中存在的对交换强度的小额比例性征税,可以防止强度永久锁定。
+
+```{code-cell} ipython3
+class Simulation:
+
+ def __init__(self, econ, seed=0):
+ self.econ = econ
+ self.rng = np.random.default_rng(seed)
+ self.agents = [Agent(i, econ, self.rng) for i in range(econ.n_types)]
+ self.types = np.repeat(np.arange(econ.n_types), econ.n_agents_per_type)
+
+ if econ.start == 'random':
+ self.holdings = self.rng.integers(0, econ.n_types, size=econ.n_agents)
+ else:
+ self.holdings = econ.produces[self.types].copy()
+ if econ.n_fiat:
+ who = self.rng.choice(econ.n_agents, size=econ.n_fiat, replace=False)
+ self.holdings[who] = econ.n_goods - 1
+
+ self.hold_hist, self.exch_hist, self.cons_hist = [], [], []
+ self.trades, self.eaten = [], []
+
+ def run(self, T, verbose=False):
+ econ, rng = self.econ, self.rng
+ evolving = econ.method != 'enumerate'
+
+ if evolving:
+ # the genetic algorithm fires on even dates with probability 1/sqrt(t/2)
+ p = 1.0 / np.sqrt(np.arange(1, T // 2 + 1))
+ even = np.arange(1, T, 2)
+ ga_dates = np.zeros(T, dtype=bool)
+ ga_dates[even] = p[:len(even)] > rng.random(len(even))
+ spec_dates = np.zeros(T, dtype=bool)
+ spec_dates[even] = p[:len(even)] > rng.random(len(even))
+
+ for t in range(1, T + 1):
+ spec = evolving and econ.method == 'ga4' and spec_dates[t - 1]
+ n_trades = n_eaten = 0
+ exch = np.zeros((econ.n_types, econ.n_goods, econ.n_goods))
+ cons = np.zeros((econ.n_types, econ.n_goods, 2))
+
+ order = rng.permutation(econ.n_agents)
+ for k in range(econ.n_agents // 2):
+ a, b = order[2 * k], order[2 * k + 1]
+ ia, ib = self.types[a], self.types[b]
+ ga, gb = self.holdings[a], self.holdings[b]
+ A, B = self.agents[ia], self.agents[ib]
+
+ # --- exchange ---
+ la, ea = A.trade_decision(ga, gb, spec)
+ lb, eb = B.trade_decision(gb, ga, spec)
+ swap = (la == 1) and (lb == 1)
+ if swap:
+ self.holdings[a], self.holdings[b] = gb, ga
+ n_trades += 1
+ exch[ia, ga, gb] += 1
+ exch[ib, gb, ga] += 1
+
+ # --- consumption ---
+ pa, pb = self.holdings[a], self.holdings[b]
+ ca, wa = A.consume_decision(pa, spec)
+ cb, wb = B.consume_decision(pb, spec)
+ cons[ia, pa, 0] += 1
+ cons[ib, pb, 0] += 1
+
+ ua = self.consume_and_produce(a, ia, pa, ca)
+ if ca == 1:
+ cons[ia, pa, 1] += 1
+ n_eaten += pa == ia
+ ub = self.consume_and_produce(b, ib, pb, cb)
+ if cb == 1:
+ cons[ib, pb, 1] += 1
+ n_eaten += pb == ib
+
+ # --- accounting ---
+ A.update(ea, wa, ua, swap or la == 0)
+ B.update(eb, wb, ub, swap or lb == 0)
+
+ if evolving:
+ if ga_dates[t - 1]:
+ self.evolve()
+ if econ.method == 'ga3':
+ for A in self.agents:
+ specialize_all(A.trade, t, rng)
+ specialize_all(A.consume, t, rng)
+ self.tax()
+
+ self.hold_hist.append(self.distribution())
+ self.exch_hist.append(exch)
+ self.cons_hist.append(cons)
+ self.trades.append(n_trades)
+ self.eaten.append(n_eaten)
+ if verbose and t % max(1, T // 10) == 0:
+ print(f" period {t:5d}: trades = {n_trades:3d},"
+ f" consumptions = {n_eaten:3d}")
+
+ def consume_and_produce(self, a, i, good, action):
+ """
+ Carry out the consumption decision and return the external payoff.
+
+ Consuming fiat money is not allowed. If the agent does consume, he
+ immediately produces his own good, so this updates his holding.
+
+ """
+ econ = self.econ
+ if action == 1 and not (econ.fiat and good == econ.n_goods - 1):
+ new = econ.produces[i]
+ self.holdings[a] = new
+ u = econ.u if good == i else 0.0
+ return u - econ.storage_costs[new]
+ return -econ.storage_costs[good]
+
+ def distribution(self):
+ "The matrix of holding frequencies pi^h_it(k)."
+ econ = self.econ
+ d = np.zeros((econ.n_types, econ.n_goods))
+ for i in range(econ.n_types):
+ h = self.holdings[self.types == i]
+ for k in range(econ.n_goods):
+ d[i, k] = np.mean(h == k)
+ return d
+
+ def tax(self):
+ if self.econ.method == 'ga4':
+ for A in self.agents:
+ T, C = A.trade, A.consume
+ P = np.where(T.action == 1, T.traded, T.used) + 1
+ T.strength -= (T.strength + 1.0) / P
+ C.strength -= (C.strength + 1.0) / (C.used + 1)
+ else:
+ for A in self.agents:
+ A.trade.strength -= 1e-4 * np.abs(A.trade.strength)
+
+ def pick_types(self):
+ "Send one type to the genetic algorithm, a second and a third each w.p. 0.33."
+ rng, n = self.rng, self.econ.n_types
+ chosen = [int(rng.integers(n))]
+ rest = [i for i in range(n) if i not in chosen]
+ while rest and len(chosen) < 3 and rng.random() < 0.33:
+ pick = rest[int(rng.integers(len(rest)))]
+ chosen.append(pick)
+ rest.remove(pick)
+ return chosen
+
+ def evolve(self):
+ econ = self.econ
+ gen = econ.method == 'ga4'
+ for i in self.pick_types():
+ genetic_algorithm(self.agents[i].trade, self.rng, generalize=gen,
+ pcross=econ.pcross, pmutation=econ.pmutation,
+ crowd_factor=8)
+ for i in self.pick_types():
+ genetic_algorithm(self.agents[i].consume, self.rng, generalize=gen,
+ pcross=econ.pcross, pmutation=econ.pmutation,
+ crowd_factor=4)
+```
+
+```{note}
+`Agent` 和 `Simulation` 引用了四个函数——`genetic_algorithm`、`specialize_all`、
+`diversify_clone` 和 `specialize_winner`——这些函数属于遗传算法,因此延后到下面的
+{ref}`mms_ga` 中介绍,在那里可以通过它们所解决的问题来阐明其动机。
+
+延后介绍它们没有任何代价。
+
+Python 在函数运行时才查找名称,而不是在定义时,我们首先研究的完全枚举经济体从不调用这四个函数中的任何一个。
+
+倾向于先看到机制再看到它的使用的读者,可以先运行该部分的代码单元格。
+```
+
+### 报告
+
+以下辅助函数将模拟输出整理成论文表格的格式。
+
+我们遵循论文的做法,报告十期移动平均值。
+
+```{code-cell} ipython3
+def good_names(econ):
+ names = [f"good {k+1}" for k in range(econ.n_types)]
+ return names + ["fiat"] if econ.fiat else names
+
+
+def type_names(econ):
+ return [f"type {i+1}" for i in range(econ.n_types)]
+
+
+def holdings(sim, t=None, window=10):
+ r"Table of $\pi^h_{it}(j)$, averaged over the `window` periods ending at `t`."
+ h = np.array(sim.hold_hist)
+ t = len(h) if t is None else t
+ d = h[max(0, t - window):t].mean(axis=0)
+ return pd.DataFrame(d, index=type_names(sim.econ),
+ columns=good_names(sim.econ)).round(3)
+
+
+def exchanges(sim, t=None, window=10):
+ r"""
+ Table of $\pi^e_{it}(jk)$: the frequency with which a type $i$ agent holds
+ good $j$, meets an agent holding good $k$, and trades. Row $j$ of column
+ $i$ holds the triple over $k$.
+ """
+ econ = sim.econ
+ e = np.array(sim.exch_hist)
+ t = len(e) if t is None else t
+ f = e[max(0, t - window):t].mean(axis=0) / econ.n_agents_per_type
+ cols = {type_names(econ)[i]:
+ ["(" + ", ".join(f"{f[i, j, k]:.2f}" for k in range(econ.n_goods)) + ")"
+ for j in range(econ.n_goods)]
+ for i in range(econ.n_types)}
+ return pd.DataFrame(cols, index=good_names(econ)).T
+
+
+def winning_actions(sim):
+ r"""
+ Table of $\tilde\pi^e_{it}(jk|j)$: the action chosen by the winning exchange
+ rule in each state, whether or not that state is ever visited.
+ """
+ econ = sim.econ
+ cols = {}
+ for i, A in enumerate(sim.agents):
+ col = []
+ for j in range(econ.n_goods):
+ acts = []
+ for k in range(econ.n_goods):
+ w, _ = auction(A.trade, np.concatenate([econ.code(j), econ.code(k)]))
+ acts.append('-' if w < 0 else str(int(A.trade.action[w])))
+ col.append("(" + ",".join(acts) + ")")
+ cols[type_names(econ)[i]] = col
+ return pd.DataFrame(cols, index=good_names(econ)).T
+
+
+def consume_actions(sim):
+ r"Table of the winning consumption action for each post-trade holding."
+ econ = sim.econ
+ cols = {}
+ for i, A in enumerate(sim.agents):
+ col = []
+ for j in range(econ.n_goods):
+ w, _ = auction(A.consume, econ.code(j))
+ col.append('-' if w < 0 else int(A.consume.action[w]))
+ cols[type_names(econ)[i]] = col
+ return pd.DataFrame(cols, index=good_names(econ)).T
+
+
+def strongest(rules, n=5):
+ "The n highest-strength classifiers in a population."
+ order = np.argsort(-rules.strength)[:n]
+ return pd.DataFrame({
+ 'classifier': [rule_string(rules.cond[i], rules.action[i]) for i in order],
+ 'strength': rules.strength[order].round(2),
+ 'times used': rules.used[order]})
+```
+
+两幅图:持有量的时间路径,对应于论文的图 6-9;以及系统所发现的交易模式图,对应于其图 2、4、9 和 11。
+
+```{code-cell} ipython3
+def plot_holdings(sim, title=None):
+ econ = sim.econ
+ h = np.array(sim.hold_hist)
+ names = good_names(econ)
+ fig, axes = plt.subplots(1, econ.n_types,
+ figsize=(3.2 * econ.n_types, 3.2), sharey=True)
+ axes = np.atleast_1d(axes)
+ for i, ax in enumerate(axes):
+ for k in range(econ.n_goods):
+ ax.plot(h[:, i, k], lw=1.0, label=names[k])
+ ax.set_title(f"type {i+1}")
+ ax.set_xlabel("$t$")
+ ax.set_ylim(-0.03, 1.03)
+ axes[0].set_ylabel(r"$\pi^h_{it}(j)$")
+ axes[-1].legend(frameon=False, fontsize=8, loc='center right')
+ if title:
+ fig.suptitle(title)
+ plt.tight_layout()
+ plt.show()
+
+
+def plot_flows(sim, window=100, cutoff=0.02, title=None):
+ """
+ One panel per type. An arrow from good j to good k means that agents of that
+ type give up j and receive k in trade; its width is the frequency with which
+ that exchange occurs over the last `window` periods.
+ """
+ econ = sim.econ
+ f = np.array(sim.exch_hist)[-window:].mean(axis=0) / econ.n_agents_per_type
+ names, n = good_names(econ), econ.n_goods
+ ang = np.pi / 2 + 2 * np.pi * np.arange(n) / n
+ xy = np.column_stack([np.cos(ang), np.sin(ang)])
+
+ fig, axes = plt.subplots(1, econ.n_types, figsize=(2.9 * econ.n_types, 3.1))
+ axes = np.atleast_1d(axes)
+ for i, ax in enumerate(axes):
+ for k in range(n):
+ ax.plot(*xy[k], 'o', ms=22, mfc='white', mec='black', zorder=2)
+ ax.annotate(names[k].replace(' ', '\n'), xy[k], ha='center',
+ va='center', fontsize=6.5, zorder=3)
+ for j in range(n):
+ for k in range(n):
+ if j == k or f[i, j, k] < cutoff:
+ continue
+ a, b = xy[j], xy[k]
+ d = b - a
+ ax.annotate("", xy=b - 0.24 * d, xytext=a + 0.24 * d, zorder=1,
+ arrowprops=dict(arrowstyle="-|>", color="C0",
+ lw=1 + 8 * f[i, j, k], alpha=0.7,
+ connectionstyle="arc3,rad=0.15"))
+ ax.set_title(f"type {i+1}", fontsize=10)
+ ax.set_xlim(-1.45, 1.45)
+ ax.set_ylim(-1.45, 1.45)
+ ax.set_aspect('equal')
+ ax.axis('off')
+ if title:
+ fig.suptitle(title)
+ plt.tight_layout()
+ plt.show()
+```
+
+## 经济体 A1.1:交换媒介会出现吗?
+
+我们的第一个经济体是维克塞尔三角,其参数为
+
+$$
+s_1 = 0.1, \quad s_2 = 1, \quad s_3 = 20, \quad u_i = 100 ,
+$$
+
+每种类型五十个主体,并对 72 个交换分类器和 12 个消费分类器进行完全枚举,全部强度为零。
+
+由于所有强度最初都相等,最初的拍卖获胜者实际上是任意的:主体一开始不知道该做什么。
+
+条件 {eq}`mms_fundcond` 在这里很容易满足,因此基本均衡是 Kiyotaki-Wright 的预测。
+
+```{code-cell} ipython3
+economy_a11 = Economy(
+ name='A1.1',
+ produces=np.array([1, 2, 0]), # type 1 -> good 2, etc. (0-indexed)
+ storage_costs=np.array([0.1, 1.0, 20.0]),
+ u=100.0,
+ method='enumerate',
+)
+
+sim_a11 = Simulation(economy_a11, seed=42)
+sim_a11.run(1000, verbose=True)
+```
+
+以下是 $t = 500$ 和 $t = 1000$ 时的持有量。
+
+```{code-cell} ipython3
+holdings(sim_a11, t=500)
+```
+
+```{code-cell} ipython3
+holdings(sim_a11)
+```
+
+将这些与上面列出的基本均衡进行比较:类型 1 以概率一持有商品 2,类型 3 以概率一持有商品 1,类型 2 在商品 1 和商品 3 之间平分。
+
+与其断言这一比较,不如让我们计算它。
+
+```{code-cell} ipython3
+fundamental = np.array([[0.0, 1.0, 0.0],
+ [0.5, 0.0, 0.5],
+ [1.0, 0.0, 0.0]])
+paper_a11 = np.array([[0.0, 1.0, 0.0], # the paper's table at t = 1000
+ [0.506, 0.0, 0.494],
+ [1.0, 0.0, 0.0]])
+
+simulated = holdings(sim_a11).to_numpy()
+print(f"max |simulated - fundamental equilibrium| = "
+ f"{np.abs(simulated - fundamental).max():.3f}")
+print(f"max |simulated - paper's table| = "
+ f"{np.abs(simulated - paper_a11).max():.3f}")
+```
+
+正如时间路径所示,收敛几乎是即时的。
+
+```{code-cell} ipython3
+plot_holdings(sim_a11, title="Economy A1.1")
+```
+
+只有类型 2 的主体保留了任何随机性,其原因很有启发性:他们是把商品 1 用作货币的主体,因此他们持有哪种商品取决于他们在“获取货币、花费货币”这一循环中所处的位置。
+
+现在让我们看看交易本身。
+
+第 $i$ 行第 $j$ 列的条目是类型 $i$ 的主体持有 $j$、遇到持有 $k$ 的人并交易的频率在 $k$ 上的三元组。
+
+```{code-cell} ipython3
+exchanges(sim_a11)
+```
+
+交易模式作为图片会更容易阅读。
+
+```{code-cell} ipython3
+plot_flows(sim_a11, title="Economy A1.1: discovered exchange pattern")
+```
+
+这正是基本均衡三角形:类型 1 放弃商品 2 换取商品 1,类型 3 放弃商品 1 换取商品 3,类型 2 完成两段流程,先放弃商品 3 换取商品 1,随后再放弃商品 1 换取商品 2。
+
+商品 1 作为货币流通。
+
+我们也可以问一问获胜规则在*从未*被访问的状态下会怎么做,这正是 Kiyotaki 和 Wright 的策略所规定的,也是论文所报告的内容。
+
+```{code-cell} ipython3
+winning_actions(sim_a11)
+```
+
+阅读“type 1”行,"good 2" 列中的条目是三元组
+$(\tilde\pi^e_1(21|2), \tilde\pi^e_1(22|2), \tilde\pi^e_1(23|2))$。
+
+一个持有商品 2 的类型 1 主体会接受商品 1,拒绝商品 3:他不进行投机。
+
+最后,我们可以深入一个分类器系统内部,读出赢得竞争的规则。
+
+```{code-cell} ipython3
+strongest(sim_a11.agents[0].trade)
+```
+
+```{code-cell} ipython3
+strongest(sim_a11.agents[0].consume)
+```
+
+对类型 1 主体而言,最强的交换规则是 `0110 -> 1`,即“如果我持有商品 2 (`01`),我的伙伴持有商品 1 (`10`),则交易”——而它被使用了数千次,而其余的对手规则则几乎从未被使用。
+
+最强的消费规则是 `10 -> 1`:“如果我持有商品 1,就吃掉它”。
+
+这正是论文所展示的、支持基本均衡的分类器 $e^1_{2,1,1}$ 和 $c^1_{1,1}$。
+
+## 经济体 A2.1:当理论预测投机时
+
+经济体 A2 与 A1 唯一的区别在于 $u_i = 500$ 而非 $100$。
+
+这足以违反 Kiyotaki-Wright 不等式 {eq}`mms_fundcond`,因此对于耐心的主体来说,此时唯一的平稳理性预期均衡变成了**投机**均衡,在其中类型 1 主体预期能迅速将商品 3 换成商品 1,因而接受商品 3。
+
+我们的适应性主体能找到它吗?
+
+```{code-cell} ipython3
+economy_a21 = Economy(
+ name='A2.1',
+ produces=np.array([1, 2, 0]),
+ storage_costs=np.array([0.1, 1.0, 20.0]),
+ u=500.0,
+ method='enumerate',
+)
+
+sim_a21 = Simulation(economy_a21, seed=42)
+sim_a21.run(1000)
+
+holdings(sim_a21)
+```
+
+以下是 Kiyotaki-Wright 理论针对这些参数所预测的投机均衡,供比较。
+
+```{code-cell} ipython3
+speculative = pd.DataFrame([[0, 0.707, 0.293],
+ [0.586, 0, 0.414],
+ [1, 0, 0]],
+ index=type_names(economy_a21),
+ columns=good_names(economy_a21))
+speculative
+```
+
+答案是否定的。
+
+模拟的持有量是*基本*均衡的持有量:类型 1 主体几乎总是持有商品 2,从不积累商品 3,而投机均衡则要求他们大约三成的时间持有商品 3。
+
+值得深入探究其原因。
+
+```{code-cell} ipython3
+winning_actions(sim_a21)
+```
+
+阅读“type 1”行、“good 2”列,获胜的交换规则拒绝商品 3:类型 1 主体不会进行投机性交换。
+
+现在看看他们如果拿到商品 3 会怎么做。
+
+```{code-cell} ipython3
+consume_actions(sim_a21)
+```
+
+```{code-cell} ipython3
+strongest(sim_a21.agents[0].consume, n=4)
+```
+
+这正是论文的诊断,从输出中可以看到。
+
+有一条消费规则 `10 -> 1` 是特定的,并且极其强:“如果我持有商品 1,就吃掉它”。
+
+决定其他一切的规则是 `## -> 1`——*不管持有什么都吃掉*——一条带两个通配符的最大程度泛化的规则,强度勉强高于零,但它却是匹配除“持有商品 1”之外任何状态的最强规则。
+
+因此,一个获得商品 3 的类型 1 主体会立即消费掉它、获得零效用,而不是把它带着并换成商品 1。
+
+正如论文所述,类型 1 主体获胜的消费分类器过于泛化——它们的 $\#$ 太多,无法区分储藏的商品——因此类型 1 主体过度消费商品 3,使得能让投机变得有利可图的信息永远无法传到交换分类器那里。
+
+作者的诊断值得引用:
+
+> *耐心需要经验。* 分类器系统内部的转移系统被设计为收敛到一组长期平均强度。在极限情况下,人工智能主体应该表现为长期平均收益最大化者……然而,最优规则要达到期望的强度需要时间。我们的人工智能主体在早期的行为可能非常短视……目前的算法似乎存在缺陷,即使在我们运行的长时间模拟中,它的实验也太少,不足以支持投机均衡。
+
+在早期,在任何规则积累起有意义的平均值之前,主体的行为是短视的,而短视的主体永远不会接受昂贵的商品。
+
+等到强度稳定下来时,其他所有人已经停止投机了,因此投机不再有利可图。
+
+被选中的均衡是*学习动态*的结果,而不仅仅是收益本身的结果。
+
+## 经济体 B.1:一种不同的生产模式
+
+经济体 B 改变了生产技术——类型 1 生产商品 3,类型 2 生产商品 1,类型 3 生产商品 2——并将储藏成本压缩为 $s = (1, 4, 9)$,$u_i = 100$。
+
+基本均衡和投机均衡都存在。
+
+论文报告了一个引人注目的现象:在 $t = 500$ 时经济体看起来像是投机均衡,但到 $t = 1000$ 时它已经转向了基本均衡。
+
+```{code-cell} ipython3
+economy_b1 = Economy(
+ name='B.1',
+ produces=np.array([2, 0, 1]), # type 1 -> good 3, type 2 -> good 1, ...
+ storage_costs=np.array([1.0, 4.0, 9.0]),
+ u=100.0,
+ method='enumerate',
+ b_trade=(0.25, 0.25),
+)
+
+sim_b1 = Simulation(economy_b1, seed=42)
+sim_b1.run(1000)
+```
+
+```{code-cell} ipython3
+holdings(sim_b1, t=500)
+```
+
+```{code-cell} ipython3
+holdings(sim_b1)
+```
+
+```{code-cell} ipython3
+plot_holdings(sim_b1, title="Economy B.1")
+```
+
+```{code-cell} ipython3
+paper_b1 = np.array([[0.0, 0.28, 0.72], # the paper's table at t = 1000
+ [0.994, 0.0, 0.006],
+ [0.526, 0.474, 0.0]])
+
+print(f"max |simulated - paper's table| = "
+ f"{np.abs(holdings(sim_b1).to_numpy() - paper_b1).max():.3f}")
+```
+
+最终状态重现了论文的定性模式,最大的单个差异约为十分之一。
+
+类型 2 主体总是持有商品 1,而类型 3 主体在商品 1 和他们自己生产的商品 2 之间分配持有量。
+
+论文报告了我们这次运行没有重现的另一个特征。
+
+在作者的模拟中,类型 3 主体在 $t = 500$ 时以概率一持有商品 2——拒绝了更便宜的商品 1,这是投机模式——直到后来才转向商品 1:
+
+> 经济体 B.1 展现出一个有趣的演化模式。在第 500 次迭代时,持有量分布,尤其是交易模式,对应于投机均衡。然而,经济体后来逐渐远离这一状态,到第 1000 次迭代时已经实际上收敛到了基本均衡。
+
+在我们的实现中,这一转变发生在最初的五十期以内,因此投机阶段在 $t = 500$ 之前就已经结束。
+
+暂态显然对实现细节非常敏感,而极限情况则并非如此。
+
+不过,作者观察到的经济学原理仍然值得说明,因为它关系到我们该如何理解经济体 A2。
+
+在那里,人们可能会怀疑基本均衡的被选中仅仅是因为短视的主体从一开始就拒绝昂贵的商品。
+
+而在经济体 B 中,系统一开始就*处于*投机模式,随着强度的积累却*远离*了它,这表明基本均衡的选择是学习动态所做出的行为,而不仅仅是初始条件的产物。
+
+```{code-cell} ipython3
+plot_flows(sim_b1, title="Economy B.1: discovered exchange pattern")
+```
+
+(mms_ga)=
+## 遗传算法
+
+完全枚举之所以可行,只是因为 Kiyotaki-Wright 的状态空间很小。
+
+在任何更大的问题中,所有可想象的规则的列表太长而无法维护,主体必须转而使用一个持续修订的有限规则种群。
+
+Marimon、McGrattan 和 Sargent 为这种情形添加了四种操作。
+
+**创造**和**多样化**我们已经实现;它们处理未预见的状态,并保证两种动作都会被尝试。
+
+对于五种商品的经济体,我们将使用一种多样化的变体,即用相反的动作克隆获胜规则,从而保留获胜者的泛化水平,而不是制造一条完全特定的规则。
+
+```{code-cell} ipython3
+def diversify_clone(rules, matches, winner, ufitness=0.5):
+ "Copy the winner with the opposite action over a rarely used rule."
+ if len(set(rules.action[matches])) > 1:
+ return
+ losers = np.flatnonzero(rules.used / (rules.used.max() + 1) < ufitness)
+ if losers.size == 0:
+ return
+ j = losers[np.argmin(rules.strength[losers])]
+ rules.replace(j, rules.cond[winner].copy(), 1 - rules.action[winner],
+ rules.strength[winner], used=rules.used[winner])
+```
+
+**特殊化**将通配符转变为特定的比特,因此一直服务于多个状态的规则可以分裂出自身更尖锐的版本。
+
+它以随时间下降的概率 $f_s(t) = 1/(2\sqrt{t})$ 被调用:实验在早期便宜,在后期则昂贵。
+
+```{code-cell} ipython3
+def specialize_all(rules, t, rng):
+ "Replace each wildcard by a bit with probability 1 / (2 sqrt(t))."
+ hit = (rules.cond == WILD) & (rng.random(rules.cond.shape)
+ < 1.0 / (2.0 * np.sqrt(t)))
+ if hit.any():
+ rules.cond[hit] = rng.integers(0, 2, size=hit.sum())
+
+
+def specialize_winner(rules, winner, pmutation, rng, ufitness=0.5):
+ "Plant a sharpened copy of a heavily used winner over a rarely used rule."
+ cond = rules.cond[winner]
+ if not np.any(cond == WILD):
+ return
+ if rules.used[winner] / (rules.used.max() + 1) <= ufitness:
+ return
+ pick = (cond == WILD) & (rng.random(cond.shape) < pmutation)
+ losers = np.flatnonzero(rules.used / (rules.used.max() + 1) < ufitness)
+ if not pick.any() or losers.size == 0:
+ return
+ j = losers[np.argmin(rules.strength[losers])]
+ new = cond.copy()
+ new[pick] = rng.integers(0, 2, size=pick.sum())
+ rules.replace(j, new, rules.action[winner], rules.strength[winner],
+ used=rules.used[winner])
+```
+
+**泛化**是真正的遗传算法。
+
+强度弱或很少被使用的规则会成为被替换的候选。
+
+父本以两个阶段抽取——首先是按照规则被使用的频率进行加权的一个子集,然后在该子集中根据强度进行轮盘赌选择——这样一条规则必须既成功又*相关*才能繁殖。
+
+两个子代通过交叉,以及在 `'ga3'` 变体中通过变异形成。
+
+然后每个子代取代它在可替换规则中最相似的那条规则,这种手段被称为*拥挤*(crowding),它通过让子代与自己的同类竞争来保持多样性。
+
+`'ga4'` 变体用一种*泛化*交叉取代了单点交叉:在一个随机抽取的区间内,父本不一致的位置会变成通配符。
+
+这正是论文第 6 节所描述、其图 5 所展示的算子,它制造出泛化的规则,而不是重组特定的规则。
+
+```{code-cell} ipython3
+def roulette(weights, rng):
+ total = weights.sum()
+ if total <= 0:
+ return int(rng.integers(0, len(weights)))
+ return int(np.searchsorted(np.cumsum(weights), rng.random() * total))
+
+
+def crowding_victim(rules, cond, action, cankill, rng, crowd_factor, crowd_subpop):
+ "De Jong crowding: the child displaces the replaceable rule it resembles most."
+ size = max(1, int(crowd_subpop * len(cankill)))
+ best, best_sim = cankill[0], -1
+ for _ in range(crowd_factor):
+ pool = (cankill if size >= len(cankill)
+ else list(rng.choice(cankill, size=size, replace=False)))
+ cand = min(pool, key=lambda i: rules.strength[i])
+ sim = (np.count_nonzero(cond == rules.cond[cand])
+ + (action != rules.action[cand]))
+ if sim > best_sim:
+ best, best_sim = cand, sim
+ return best
+
+
+def genetic_algorithm(rules, rng, generalize=False, pcross=0.6, pmutation=0.01,
+ propselect=0.2, propused=0.7, crowd_factor=8,
+ crowd_subpop=0.5, uratio=(0.0, 0.2)):
+ n, L = rules.n, rules.length
+ if n < 4:
+ return
+
+ # rules that are weak or seldom used may be replaced
+ max_used = max(rules.used.max() + (1 if generalize else 0), 1)
+ cankill = list(np.flatnonzero((rules.strength < uratio[0]) |
+ (rules.used / max_used < uratio[1])))
+ if not cankill:
+ return
+
+ fitness = rules.strength - min(rules.strength.min(), 0.0) + 1e-6
+ n_pairs = min(max(1, round(propselect * n * 0.5)), (len(cankill) + 1) // 2)
+ n_called = int(propused * n)
+
+ for _ in range(n_pairs):
+ if not cankill:
+ break
+
+ # stage 1: pre-select a pool with probability proportional to usage
+ if n_called < n:
+ avail = rules.used.astype(float) + 1.0
+ pool = []
+ for _ in range(n_called):
+ if avail.sum() <= 0:
+ break
+ k = roulette(avail, rng)
+ pool.append(k)
+ avail[k] = 0.0
+ if len(pool) < 2:
+ pool = list(range(n))
+ else:
+ pool = list(range(n))
+ pool = np.array(pool)
+
+ # stage 2: roulette wheel on fitness within the pool
+ mum, dad = pool[roulette(fitness[pool], rng)], pool[roulette(fitness[pool], rng)]
+ kids = [rules.cond[mum].copy(), rules.cond[dad].copy()]
+ acts = [rules.action[mum], rules.action[dad]]
+ avg = 0.5 * (rules.strength[mum] + rules.strength[dad])
+
+ if generalize:
+ # two-point crossover in which disagreements become wildcards
+ lo, hi = np.sort(rng.integers(0, L + 1, size=2))
+ inside = rng.random() > 0.5
+ region = (np.arange(lo, hi) if inside else
+ np.concatenate([np.arange(0, lo), np.arange(hi, L)]))
+ for j in region:
+ a, b = kids[0][j], kids[1][j]
+ if a >= 0 and b >= 0 and a != b:
+ kids[0][j] = kids[1][j] = WILD
+ else:
+ # single-point crossover with ternary mutation
+ jc = 1 + int((L - 1) * rng.random()) if rng.random() < pcross else L
+ kids = [np.concatenate([rules.cond[mum][:jc], rules.cond[dad][jc:]]),
+ np.concatenate([rules.cond[dad][:jc], rules.cond[mum][jc:]])]
+ for k in range(2):
+ flip = rng.random(L) < pmutation
+ if flip.any():
+ shift = rng.integers(1, 3, size=flip.sum())
+ kids[k][flip] = ((kids[k][flip] + 1 + shift) % 3) - 1
+ if rng.random() < pmutation:
+ acts[k] = 1 - acts[k]
+
+ for k, parent in zip(range(2), (mum, dad)):
+ if not cankill:
+ break
+ j = crowding_victim(rules, kids[k], acts[k], cankill, rng,
+ crowd_factor, crowd_subpop)
+ rules.replace(j, kids[k], acts[k], avg,
+ used=rules.used[parent], traded=rules.traded[parent])
+ cankill.remove(j)
+```
+
+## 经济体 A1.2:从随机规则中学习
+
+经济体 A1.2 的参数与 A1.1 相同,但主体现在从随机抽取的 72 条交换规则和 12 条消费规则开始。
+
+这些规则中的大多数都是无意义的,也不能保证种群中甚至包含支持某个均衡所需的规则。
+
+遗传算子必须制造出这些规则。
+
+```{code-cell} ipython3
+economy_a12 = Economy(
+ name='A1.2',
+ produces=np.array([1, 2, 0]),
+ storage_costs=np.array([0.1, 1.0, 20.0]),
+ u=100.0,
+ method='ga3',
+)
+
+sim_a12 = Simulation(economy_a12, seed=2)
+sim_a12.run(2000, verbose=True)
+```
+
+```{code-cell} ipython3
+holdings(sim_a12, t=1000)
+```
+
+```{code-cell} ipython3
+holdings(sim_a12)
+```
+
+```{code-cell} ipython3
+plot_holdings(sim_a12, title="Economy A1.2")
+```
+
+再次达到了基本均衡,尽管这需要花费明显更长的时间,且早期的转变过程比完全枚举时要杂乱得多。
+
+让我们看看哪些规则存活了下来。
+
+```{code-cell} ipython3
+strongest(sim_a12.agents[0].trade, n=6)
+```
+
+种群已经收敛到几乎完全一样的单一规则的复制品上。
+
+`0110 -> 1`——“持有商品 2,遇到商品 1,交易”——正是完全枚举在经济体 A1.1 中所选出的规则,但在这里遗传算法不得不构建出它,现在种群中有好几份该规则的副本。
+
+其他条目则是它的后代:`0111 -> 1` 有相同的自身持有条件,但伙伴条件为 `11`,这不是任何一种商品的编码,因此该规则永远无法触发。
+
+它之所以携带高强度和大使用计数,只是因为子代同时继承了父代的强度和使用计数。
+
+```{code-cell} ipython3
+plot_flows(sim_a12, title="Economy A1.2: discovered exchange pattern")
+```
+
+```{warning}
+收敛并不能保证每次运行都会发生。
+
+在随机初始规则的情况下,种群有时会未能在 2000 期内制造出支持基本均衡所需的规则,经济体就会陷入几乎没有交易的模式。
+
+尝试更改上面的种子来观察这一点。
+
+论文明确指出,该算法的“实验太少”,改进它是未完成的工作。
+```
+
+## 经济体 C:法定货币
+
+现在添加第四种物品,商品 0,它满足
+
+* 储藏无成本,$s_0 = 0$;
+* 对任何人都不产生效用;且
+* 不能被消费。
+
+它本质上是毫无价值的。
+
+它是通过在 $t = 0$ 时将 48 个单位交给 48 个随机选出的主体来引入的,商品的储藏成本被提高到 $s = (9, 14, 29)$,使得没有一种商品的储藏成本能接近货币那么低。
+
+如果这样的商品流通起来,那只能是因为主体已经发现其他主体会接受它。
+
+```{code-cell} ipython3
+economy_c = Economy(
+ name='C',
+ produces=np.array([1, 2, 0]),
+ storage_costs=np.array([9.0, 14.0, 29.0, 0.0]), # last good is fiat money
+ u=100.0,
+ method='ga3',
+ n_trade_rules=150,
+ n_consume_rules=20,
+ b_consume=(0.025, 0.25),
+ n_fiat=48,
+ start='production',
+)
+
+sim_c = Simulation(economy_c, seed=2)
+sim_c.run(2000, verbose=True)
+```
+
+```{code-cell} ipython3
+holdings(sim_c, t=750)
+```
+
+```{code-cell} ipython3
+holdings(sim_c, t=1250)
+```
+
+与论文在 $t = 1250$ 时的表格进行比较,其中各列为(商品 1、商品 2、商品 3、法定货币):类型 1 为 $(0, 0.54, 0, 0.46)$,类型 2 为 $(0.18, 0, 0.53, 0.28)$,类型 3 为 $(0.77, 0, 0, 0.21)$。
+
+每种类型都在相当一部分时间里持有法定货币。
+
+```{code-cell} ipython3
+plot_holdings(sim_c, title="Economy C: fiat money")
+```
+
+```{code-cell} ipython3
+plot_flows(sim_c, title="Economy C: discovered exchange pattern")
+```
+
+```{code-cell} ipython3
+winning_actions(sim_c)
+```
+
+进出法定货币节点的箭头显示了每种类型都在双向进行货币的转手:主体放弃商品以获取货币,也放弃货币以获取他们消费的商品。
+
+环境中没有任何东西告诉他们这样做。
+
+每个主体只是发现,最终持有那种无成本物品的规则比不这样做的规则赚得更多——而由于所有人都在同时发现这一点,这种信念变得自我实现。
+
+这是一种真正的社会性安排,由那些个体上对它一无所知的主体所构建。
+
+## 经济体 D:五种商品,五种类型
+
+最后一个经济体有五种类型和五种商品,生产模式为
+
+| 类型 $i$ | 生产 | 消费 |
+|---|---|---|
+| 1 | 商品 3 | 商品 1 |
+| 2 | 商品 4 | 商品 2 |
+| 3 | 商品 5 | 商品 3 |
+| 4 | 商品 1 | 商品 4 |
+| 5 | 商品 2 | 商品 5 |
+
+储藏成本 $s = (1, 4, 9, 16, 30)$,$u_i = 200$。
+
+商品现在用三位编码,每个主体携带 180 条交换规则和 20 条消费规则——这只是可能规则中的一小部分,因此遗传算法是必不可少的。
+
+作者强调,在运行之前,他们对这个经济体*没有解析的均衡刻画*。
+
+模拟被用作一种*发现*均衡可能是什么样子的装置,之后可以对其进行解析验证。
+
+```{code-cell} ipython3
+economy_d = Economy(
+ name='D',
+ produces=np.array([2, 3, 4, 0, 1]), # type 1 -> good 3, type 2 -> good 4, ...
+ storage_costs=np.array([1.0, 4.0, 9.0, 16.0, 30.0]),
+ u=200.0,
+ method='ga4',
+ n_trade_rules=180,
+ n_consume_rules=20,
+ start='production',
+)
+
+sim_d = Simulation(economy_d, seed=3)
+sim_d.run(2000, verbose=True)
+```
+
+```{code-cell} ipython3
+holdings(sim_d, t=500)
+```
+
+```{code-cell} ipython3
+holdings(sim_d)
+```
+
+```{code-cell} ipython3
+plot_holdings(sim_d, title="Economy D: five goods, five types")
+```
+
+有两个特征引人注目。
+
+第一,对角线为零:没有一种类型会在一期结束时持有它所消费的商品,因为它把它消费掉了。
+
+第二,每种类型都会积累自己的生产商品,以及比它更*便宜*储藏的商品,从不积累更昂贵的商品。
+
+例如,类型 3 生产商品 5,即经济体中最昂贵的商品,只有部分时间持有它,因为已经将其中一部分换成了便宜得多的商品 2。
+
+类型 4 生产商品 1,即所有商品中最便宜的,就干脆持有它。
+
+让我们对实际发生的交易进行分类。
+
+```{code-cell} ipython3
+def trade_composition(sim, window=200):
+ """
+ Classify realized trades by what the agent acquires: its own consumption
+ good, a good that is cheaper to store than the one given up, or neither.
+ """
+ econ = sim.econ
+ f = np.array(sim.exch_hist)[-window:].sum(axis=0)
+ s = econ.storage_costs
+ counts = {'own consumption good': 0.0, 'a cheaper good': 0.0, 'neither': 0.0}
+ for i in range(econ.n_types):
+ for j in range(econ.n_goods):
+ for k in range(econ.n_goods):
+ if j == k:
+ continue
+ key = ('own consumption good' if k == i else
+ 'a cheaper good' if s[k] < s[j] else 'neither')
+ counts[key] += f[i, j, k]
+ total = sum(counts.values())
+ return pd.Series({k: round(v / total, 3) for k, v in counts.items()},
+ name='share of realized trades')
+
+
+trade_composition(sim_d)
+```
+
+大约三分之二的实际交易获取的要么是主体自己的消费商品,要么是储藏成本更便宜的商品,这与论文所描述的模式一致。
+
+剩下的三分之一值得评论,因为它并非反对该模式的证据。
+
+考虑一个类型 1 主体,他生产商品 3,用其换取贵得多的商品 5。
+
+然后他消费商品 5——从中获得零效用——再次生产商品 3,因此他在期末仍持有商品 3、支付 $s_3$,与他拒绝交易的情形完全相同。
+
+这类交易在收益上是中性的,因此记账系统中没有任何东西会把产生这些交易的规则挤出种群。
+
+```{code-cell} ipython3
+plot_flows(sim_d, window=200, cutoff=0.03,
+ title="Economy D: discovered exchange pattern")
+```
+
+论文对这些模式所做的总结是:
+
+> 从模拟结果可以看到,交易模式几乎描绘出一个基本均衡的样子,主体只愿意为储藏成本低于当前所储藏商品的商品进行交易,除非它们总是接受本类型的商品。可以检测到一些投机性的举动。
+
+一个这样的投机性举动的例子是,类型 2 主体接受商品 3 换取商品 1,并非因为商品 3 便宜,而是因为类型 3 主体会用商品 2 来换取它。
+
+## 结束语
+
+Marimon、McGrattan 和 Sargent 的主体知道得很少。
+
+他们不知道自己的效用函数,不知道储藏成本,不知道商品在人群中的分布,当然也不求解动态规划问题。
+
+他们只是在体验效用时认出它,在承担成本时认出它,并保留滚动平均值。
+
+从中产生了以下结果:
+
+* **纳什-马尔可夫行为是可学习的。** 在所模拟的大多数经济体中,持有量和交易模式都收敛到 Kiyotaki-Wright 模型的一个平稳纳什均衡。
+* **学习在均衡之间做出选择。** 在理性预期模型同时容许基本均衡和投机均衡的地方,分类器系统总是找到基本均衡。
+ - 经济体 A2 表明,即使理论指出对于耐心的主体来说,该均衡本不该成立,它们仍能找到它。
+ - 经济体 B 表明这并非仅仅是早期的短视,因为该经济体随着学习*远离*了投机模式。
+* **制度是可以被发现的。** 经济体 C 的主体仅凭他们自己对储藏成本的体验,就建立起了一套法定货币体系。
+* **该方法具有可扩展性。** 经济体 D 在其作者尚未求解的模型中产生了一个可信的均衡描述。
+
+论文对缺失之处很坦诚。
+
+没有收敛定理,只有对随机逼近论证如何能够提供收敛定理的一个概述;作者判断自己的遗传算法提供的“实验太少”——而这正是投机均衡从未出现的原因。
+
+这一诊断——即适应性系统对均衡的选择是由它探索的多少和何时探索所支配的——已被证明是经久不衰的。
+
+## 练习
+
+```{exercise-start}
+:label: mms_ex1
+```
+
+经济体 A2.2 就是经济体 A2——$u_i = 500$——但从随机生成的规则开始,并用 `'ga3'` 遗传算法进行演化,正如 A1.2 之于 A1.1 的关系一样。
+
+论文的总结表将其均衡类型列为投机型,作者报告说在 1000 次迭代之后经济体尚未收敛,$t = 1000$ 时的交易模式比 $t = 500$ 时更接近基本均衡。
+
+将其模拟 2000 期,并报告 $t = 500$、$t = 1000$ 和 $t = 2000$ 时的持有量。
+
+遗传算法提供的额外实验是否会产生投机?
+
+```{exercise-end}
+```
+
+```{solution-start} mms_ex1
+:class: dropdown
+```
+
+```{code-cell} ipython3
+economy_a22 = Economy(
+ name='A2.2',
+ produces=np.array([1, 2, 0]),
+ storage_costs=np.array([0.1, 1.0, 20.0]),
+ u=500.0,
+ method='ga3',
+)
+
+sim_a22 = Simulation(economy_a22, seed=3)
+sim_a22.run(2000)
+
+for t in (500, 1000, 2000):
+ print(f"\nt = {t}")
+ print(holdings(sim_a22, t=t))
+```
+
+```{code-cell} ipython3
+plot_holdings(sim_a22, title="Economy A2.2")
+```
+
+经济体再次落定在基本模式上:类型 1 持有商品 2,类型 3 持有商品 1,类型 2 在商品 1 和商品 3 之间交替。
+
+类型 1 主体并未持有商品 3,因此他们没有进行投机。
+
+将初始规则随机化并让遗传算法运行本身并不能产生足够的实验来维持投机——这正是论文自己的结论。
+
+```{solution-end}
+```
+
+```{exercise-start}
+:label: mms_ex2
+```
+
+经济体 B.2 是从随机规则出发、使用 `'ga3'` 遗传算法的经济体 B。
+
+论文报告说它“在 2000 期后仍未收敛”,但正“朝着基本均衡移动”,在 $t = 2000$ 时的持有量为:类型 1 为 $(0, 0.354, 0.646)$,类型 2 为 $(0.996, 0, 0.004)$,类型 3 为 $(0.268, 0.732, 0)$。
+
+对其进行模拟,并与我们上面运行的完全枚举经济体 B.1 进行比较。
+
+```{exercise-end}
+```
+
+```{solution-start} mms_ex2
+:class: dropdown
+```
+
+```{code-cell} ipython3
+economy_b2 = Economy(
+ name='B.2',
+ produces=np.array([2, 0, 1]),
+ storage_costs=np.array([1.0, 4.0, 9.0]),
+ u=100.0,
+ method='ga3',
+)
+
+sim_b2 = Simulation(economy_b2, seed=1)
+sim_b2.run(2000)
+
+for t in (500, 1000, 2000):
+ print(f"B.2 at t = {t}")
+ print(holdings(sim_b2, t=t))
+ print()
+print("B.1 (complete enumeration) at t = 1000, for comparison")
+print(holdings(sim_b1))
+```
+
+```{code-cell} ipython3
+plot_holdings(sim_b2, title="Economy B.2")
+```
+
+跟踪那两种发生变化的类型。
+
+类型 2 主体在 $t = 500$ 时仍然是分裂的,到 $t = 1000$ 时已经锁定在商品 1 上。
+
+类型 3 主体一开始主要持有他们自己的生产商品 2,在同一区间内将相当一部分持有量转移到了更便宜的商品 1 上,这正是论文所描述的朝向基本均衡的移动。
+
+在那之后,类型 3 的分裂比例不再朝一个方向移动,而是从一次读数到下一次读数在半分状态附近徘徊,因此这个经济体没有像 B.1 那样收敛。
+
+这正是论文对它自己的判断。
+
+与完全枚举的运行相比,遗传算法使类型 1 和类型 2 达到了大致相同的位置,但类型 3 的进展要落后得多。
+
+从随机材料中构建所需的规则要花费时间,而枚举经济体从不需要花费这些时间。
+
+```{solution-end}
+```
+
+```{exercise-start}
+:label: mms_ex3
+```
+
+出价函数 $b_1(e) = b_{11} + b_{12}\sigma_e$ 偏向特定规则而非泛化规则,因为 $\sigma_e$ 随通配符数量的增加而下降。
+
+如果去掉这种倾向会发生什么?
+
+将 $b_{12} = 0$,保持 $b_{11} + b_{12}$ 固定为 $0.05$,重新运行经济体 A1.1,并将获胜规则中的通配符数量与基准情形进行比较。
+
+单次运行不足以解决这个问题,因此要在多个种子上取平均值。
+
+```{exercise-end}
+```
+
+```{solution-start} mms_ex3
+:class: dropdown
+```
+
+```{code-cell} ipython3
+economy_flat = Economy(
+ name='A1.1 with no specificity premium',
+ produces=np.array([1, 2, 0]),
+ storage_costs=np.array([0.1, 1.0, 20.0]),
+ u=100.0,
+ method='enumerate',
+ b_trade=(0.05, 0.0),
+)
+
+sim_flat = Simulation(economy_flat, seed=42)
+sim_flat.run(1000)
+
+print(holdings(sim_flat))
+```
+
+持有量没有变化:无论哪种方式都能达到基本均衡。
+
+发生变化的是*到达那里的规则种类*。
+
+```{code-cell} ipython3
+def wildcards_in_winners(sim):
+ "Average number of wildcards in the exchange rule that wins in each state."
+ econ = sim.econ
+ counts = []
+ for A in sim.agents:
+ for j in range(econ.n_goods):
+ for k in range(econ.n_goods):
+ w, _ = auction(A.trade, np.concatenate([econ.code(j), econ.code(k)]))
+ if w >= 0:
+ counts.append(np.count_nonzero(A.trade.cond[w] == WILD))
+ return np.mean(counts)
+
+
+def average_wildcards(b_trade, seeds=range(8)):
+ out = []
+ for seed in seeds:
+ econ = Economy(name='sweep', produces=np.array([1, 2, 0]),
+ storage_costs=np.array([0.1, 1.0, 20.0]), u=100.0,
+ method='enumerate', b_trade=b_trade)
+ sim = Simulation(econ, seed=seed)
+ sim.run(1000)
+ out.append(wildcards_in_winners(sim))
+ return np.array(out)
+
+
+base = average_wildcards((0.025, 0.025))
+flat = average_wildcards((0.05, 0.0))
+
+pd.DataFrame({"baseline $b_{12} = 0.025$": base.round(2),
+ "no premium $b_{12} = 0$": flat.round(2)},
+ index=pd.Index(range(8), name="seed"))
+```
+
+```{code-cell} ipython3
+print(f"mean wildcards, baseline : {base.mean():.3f}")
+print(f"mean wildcards, no specificity bid : {flat.mean():.3f}")
+print(f"higher without the premium in {np.sum(flat > base)} of {len(base)} seeds")
+```
+
+去掉特殊性溢价会提高获胜规则的平均泛化程度,但这一效应并不明显,也并非每次运行都会出现。
+
+这一点值得了解,而不是被一带而过。
+
+在完全枚举的情形下,特定规则从一开始就已全部存在并自行积累强度,因此出价溢价只是使拍卖倾向于它们的诸多力量之一;这种反作用是在平均值中可见的一种真实倾向,而不是每次运行都必然遵守的定律。
+
+这种倾向在更难被察觉的地方反而影响更大。
+
+正是论文归咎于经济体 A2 失败的机制:泛化的消费分类器无法区分储藏的商品,因此它们无法向交换分类器传递能使投机变得有利可图的信息。
+
+```{solution-end}
+```
\ No newline at end of file
diff --git a/lectures/phillips_adaptive.md b/lectures/phillips_adaptive.md
index 09dd8b3..5c68b17 100644
--- a/lectures/phillips_adaptive.md
+++ b/lectures/phillips_adaptive.md
@@ -35,6 +35,9 @@ translation:
# 适应性预期与费尔普斯问题
+```{index} single: Phillips Curve; Adaptive Expectations
+```
+
```{contents} Contents
:depth: 2
```
@@ -60,6 +63,10 @@ translation:
本讲座遵循 {cite}`Sargent1999` 第 5 章的内容。
+{doc}`phillips_credible_policies` 追求优于纳什结果的结果,同时保持*每个人*都是理性的,结果发现了太多这样的结果:一个可持续值的连续统,而理论内部没有任何依据可以从中做出选择。
+
+在这里,我们以尽可能小的方式退让这种完美性,保持政府是理性的,但给予公众一个机械式的预测规则。
+
我们将描述:
* 凯根-弗里德曼适应性预期假设,
@@ -115,6 +122,12 @@ x_t = (1 - \lambda) \sum_{i=1}^{\infty} \lambda^{i-1} y_{t-i} .
让我们用数值方法验证这个归纳性质。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "适应性预期收敛到恒定通货膨胀政策:归纳性质"
+ name: fig-adapt-induction
+---
def adaptive_forecast(y, λ, x0=0.0):
"Simulate x_t = λ x_{t-1} + (1-λ) y_{t-1}."
T = len(y)
@@ -168,26 +181,26 @@ U_t = U^* - \theta(y_t - x_t) .
因此,政府的问题是一个带有以下要素的贴现 LQ 控制问题:
-* 状态 $s_t = \begin{bmatrix} 1 & x_t \end{bmatrix}'$,
+* 状态 $s_t = \begin{bmatrix} 1 & x_t \end{bmatrix}^\top$,
* 控制 $y_t$,以及
* 转移方程 $x_{t+1} = \lambda x_t + (1 - \lambda) y_t$。
### 将问题转化为 LQ 形式
-令 $U_t = a' s_t - \theta y_t$,其中 $a = \begin{bmatrix} U^* & \theta \end{bmatrix}'$。
+令 $U_t = a^\top s_t - \theta y_t$,其中 $a = \begin{bmatrix} U^* & \theta \end{bmatrix}^\top$。
那么每期损失 $\tfrac{1}{2}(U_t^2 + y_t^2)$ 等于
$$
-\frac{1}{2}\left[ s_t'(a a') s_t + (\theta^2 + 1) y_t^2 - 2 \theta \, y_t \, (a' s_t) \right] .
+\frac{1}{2}\left[ s_t^\top (a a^\top) s_t + (\theta^2 + 1) y_t^2 - 2 \theta \, y_t \, (a^\top s_t) \right] .
$$
-将其与 QuantEcon 的 LQ 损失 $s_t' R s_t + y_t' Q y_t + 2 y_t' N s_t$ 进行匹配,并将转移方程与 $s_{t+1} = A s_t + B y_t$ 进行匹配,得到
+将其与 QuantEcon 的 LQ 损失 $s_t^\top R s_t + y_t^\top Q y_t + 2 y_t^\top N s_t$ 进行匹配,并将转移方程与 $s_{t+1} = A s_t + B y_t$ 进行匹配,得到
$$
-R = \tfrac{1}{2} a a', \quad
+R = \tfrac{1}{2} a a^\top, \quad
Q = \tfrac{1}{2}(\theta^2 + 1), \quad
-N = -\tfrac{1}{2}\theta\, a', \quad
+N = -\tfrac{1}{2}\theta\, a^\top, \quad
A = \begin{bmatrix} 1 & 0 \\ 0 & \lambda \end{bmatrix}, \quad
B = \begin{bmatrix} 0 \\ 1 - \lambda \end{bmatrix} .
$$
@@ -290,6 +303,12 @@ for δ in [0.96, 1.0]:
让我们绘制完整的反通货膨胀路径。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 菲尔普斯问题在有贴现和无贴现情形下的最优反通货膨胀路径
+ name: fig-adapt-disinflation
+---
fig, axes = plt.subplots(1, 2, figsize=(11, 4.5))
for δ, ax in zip([0.96, 1.0], axes):
@@ -316,23 +335,23 @@ plt.show()
定义向量
$$
-X_{U,t} = \begin{bmatrix} U_{t-1} & \cdots & U_{t-m_U} \end{bmatrix}',
+X_{U,t} = \begin{bmatrix} U_{t-1} & \cdots & U_{t-m_U} \end{bmatrix}^\top,
\qquad
-X_{y,t} = \begin{bmatrix} y_{t-1} & \cdots & y_{t-m_y} \end{bmatrix}',
+X_{y,t} = \begin{bmatrix} y_{t-1} & \cdots & y_{t-m_y} \end{bmatrix}^\top,
$$
-以及状态向量 $X_t = \begin{bmatrix} X_{U,t}' & X_{y,t}' & 1 \end{bmatrix}'$,其中收集了 $t-1$ 期及更早时期的信息。
+以及状态向量 $X_t = \begin{bmatrix} X_{U,t}^\top & X_{y,t}^\top & 1 \end{bmatrix}^\top$,其中收集了 $t-1$ 期及更早时期的信息。
我们可以写出两条简化形式的菲利普斯曲线,它们仅在*拟合方向*上有所不同:
$$
-\text{古典型:} \quad U_t = \gamma' X_{C,t} + \varepsilon_{C,t},
-\qquad X_{C,t} = \begin{bmatrix} y_t & X_{t-1}' \end{bmatrix}',
+\text{古典型:} \quad U_t = \gamma^\top X_{C,t} + \varepsilon_{C,t},
+\qquad X_{C,t} = \begin{bmatrix} y_t & X_{t-1}^\top \end{bmatrix}^\top,
$$
$$
-\text{凯恩斯型:} \quad y_t = \beta' X_{K,t} + \varepsilon_{K,t},
-\qquad X_{K,t} = \begin{bmatrix} U_t & X_{t-1}' \end{bmatrix}' .
+\text{凯恩斯型:} \quad y_t = \beta^\top X_{K,t} + \varepsilon_{K,t},
+\qquad X_{K,t} = \begin{bmatrix} U_t & X_{t-1}^\top \end{bmatrix}^\top.
$$
下标 $C$ 和 $K$ 分别代表*古典型*(将 $U$ 对 $y$ 回归)和*凯恩斯型*(将 $y$ 对 $U$ 回归)。
@@ -354,7 +373,10 @@ h = h(\gamma) .
这些对象——$\gamma$、$\beta$、$h(\gamma)$,以及两种拟合方向——正是我们在 {doc}`phillips_self_confirming` 中定义**自我确认均衡**所需要的关键要素。
```{note}
-**归纳假设**是这样一种限制:在凯恩斯型菲利普斯曲线 $y_t = \beta' X_{K,t} + \varepsilon_{K,t}$ 中,滞后 $y$ 值上的权重之和为一(等价地,在古典形式中,当期与滞后 $y$ 值上的权重之和为零)。在适应性预期下,这一点成立,因为 {eq}`pa_geom` 中的权重之和为一。
+**归纳假设**是这样一种限制:在凯恩斯型菲利普斯曲线 $y_t = \beta^\top X_{K,t} + \varepsilon_{K,t}$ 中,
+滞后 $y$ 值上的权重之和为一(等价地,在古典形式中,当期与滞后 $y$ 值上的权重之和为零)。
+
+在适应性预期下,这一点成立,因为 {eq}`pa_geom` 中的权重之和为一。
```
## 检验自然率假说
@@ -430,9 +452,10 @@ for δ in δ_grid:
y_inf.append(y[-1])
fig, ax = plt.subplots(figsize=(8, 4.5))
-ax.plot(δ_grid, y_inf, 'o-')
+ax.plot(δ_grid, y_inf, 'o-', lw=2)
ax.set_xlabel(r'discount factor $\delta$')
ax.set_ylabel(r'limiting inflation $y_\infty$')
+ax.set_title('Limiting inflation falls as patience rises')
plt.show()
```
@@ -473,4 +496,4 @@ print(f"gap = {y[-1] - x[-1]:.2e}")
在稳态下,预期通货膨胀率与实际通货膨胀率一致,这证实了归纳性质使得公众的预测在极限情形下是正确的。
```{solution-end}
-```
\ No newline at end of file
+```
diff --git a/lectures/phillips_credibility.md b/lectures/phillips_credibility.md
index 2acf39b..7d04a4b 100644
--- a/lectures/phillips_credibility.md
+++ b/lectures/phillips_credibility.md
@@ -34,6 +34,9 @@ translation:
# 信誉问题
+```{index} single: Phillips Curve; Credibility Problem
+```
+
```{contents} Contents
:depth: 2
```
@@ -42,7 +45,7 @@ translation:
本讲座描述了一个由 {cite}`KydlandPrescott1977` 以及罗伯特·巴罗和大卫·戈登所研究的那种基本预期菲利普斯曲线模型。
-这是基于 {cite}`Sargent1999` 各章内容的系列讲座中的第一讲。
+这是基于 {cite}`Sargent1999` 各章内容的系列讲座中第一讲*建模*讲座,紧接 {doc}`phillips_two_stories` 之后,后者阐述了两种叙事并回顾了卢卡斯批判。
那些章节形式化了:
@@ -119,7 +122,9 @@ r(x, y) = - \frac{1}{2} \left[ \left(U^* - \theta (y - x)\right)^2 + y^2 \right]
**纳什均衡:** 满足 (i) $x = y$,以及 (ii) $y = B(x)$ 的二元组 $(x, y)$。
-**拉姆齐问题:** $\max_y r(y, y)$。*拉姆齐结果*是达到最大值的 $y$。
+**拉姆齐问题:** $\max_y r(y, y)$。
+
+*拉姆齐结果*是达到最大值的 $y$。
**最优反应动态:** 动态系统 $y_t = B(y_{t-1})$,给定 $y_0$。
@@ -218,6 +223,12 @@ print(f"Ramsey payoff r = {cm.r(y_R, y_R):.3f}")
给定 $x$ 时,政府对 $y$ 的最优反应出现在无差异曲线与由 $x$ 索引的菲利普斯曲线相切的位置。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 以预期通货膨胀率为索引的菲利普斯曲线、政府无差异曲线以及纳什和拉姆齐结果
+ name: fig-cred-nash-ramsey
+---
fig, ax = plt.subplots(figsize=(7, 6))
U_grid = np.linspace(0, 12, 200)
@@ -230,7 +241,7 @@ for x in [0.0, y_N / 2, y_N]:
# government indifference curves (circles U^2 + y^2 = const)
ξ = np.linspace(0, 2 * np.pi, 200)
-for R in [y_R, np.hypot(cm.U_star, y_N)]:
+for R in [np.hypot(cm.U_star, y_R), np.hypot(cm.U_star, y_N)]:
if R > 0:
ax.plot(R * np.cos(ξ), R * np.sin(ξ), 'C1--', lw=1)
@@ -277,6 +288,12 @@ y_path = best_response_path(cm, y0=0.0, T=20)
```
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 最优反应动态蛛网图收敛至纳什通货膨胀率
+ name: fig-cred-best-response
+---
fig, ax = plt.subplots(figsize=(6, 6))
x_grid = np.linspace(0, y_N + 1, 100)
@@ -374,12 +391,18 @@ def ls_learning(cm, T=2000, σ_η=1.0, seed=0):
for t in range(1, T + 1):
η = σ_η * rng.standard_normal()
y[t] = cm.B(x[t - 1]) + η
- gain = 1.0 / (t + 1) # decreasing gain
+ gain = 1.0 / t # the 1/(t-1) gain of equation (7), reindexed
x[t] = x[t - 1] + gain * (cm.B(x[t - 1]) - x[t - 1] + η)
return x, y
```
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 预期通货膨胀率的最小二乘学习收敛于纳什结果
+ name: fig-cred-ls-learning
+---
x, y = ls_learning(cm, T=2000, σ_η=1.0)
fig, ax = plt.subplots(figsize=(9, 5))
@@ -410,9 +433,11 @@ plt.show()
后续讲座描述了三种建模前瞻性的方式,它们赋予不同程度的理性,并预测不同质量的结果:
-1. 一种声誉方法,将理性预期同时赋予政府和公众。许多结果都是可持续的,从拉姆齐结果的重复到比纳什结果重复更差的路径都有可能。
-2. 一种方法,保持政府的理性,但赋予公众原始凯根-弗里德曼意义上的*适应性*预期。这是 {doc}`phillips_adaptive` 的主题。根据贴现因子与适应参数的比较,这种设定可以改善结果,甚至可能维持拉姆齐结果的重复。
-3. 一种将适应性行为同时赋予政府和公众的方法。这是 {doc}`phillips_misspecified` 和 {doc}`phillips_self_confirming` 的主题。
+1. 一种声誉方法,将理性预期同时赋予政府和公众,这是 {doc}`phillips_credible_policies` 的主题。
+ - 许多结果都是可持续的,从拉姆齐结果的重复到比纳什结果重复更差的路径都有可能,而这种多重性正是该理论的主要启示。
+2. 一种方法,保持政府的理性,但赋予公众原始凯根-弗里德曼意义上的*适应性*预期,这是 {doc}`phillips_adaptive` 的主题。
+ - 根据贴现因子与适应参数的比较,这种设定可以改善结果,甚至可能维持拉姆齐结果的重复。
+3. 一种将适应性行为同时赋予政府和公众的方法,这是 {doc}`phillips_misspecified` 和 {doc}`phillips_self_confirming` 的主题。
## 附录:随机逼近
@@ -498,6 +523,7 @@ for θ in [0.5, 1.0, 2.0]:
ax.axhline(0, color='k', lw=0.8)
ax.set_xlabel('$t$')
ax.set_ylabel('$x_t - \\theta U^*$')
+ax.set_title('按菲利普斯斜率划分的最小二乘学习')
ax.legend()
plt.show()
```
@@ -507,4 +533,4 @@ plt.show()
较大的 $\theta$ 则使 $x_t$ 在其极限附近更持久地波动。
```{solution-end}
-```
\ No newline at end of file
+```
diff --git a/lectures/phillips_credible_policies.md b/lectures/phillips_credible_policies.md
new file mode 100644
index 0000000..5dc6b5d
--- /dev/null
+++ b/lectures/phillips_credible_policies.md
@@ -0,0 +1,1017 @@
+---
+jupytext:
+ text_representation:
+ extension: .md
+ format_name: myst
+ format_version: 0.13
+ jupytext_version: 1.16.7
+kernelspec:
+ display_name: Python 3 (ipykernel)
+ language: python
+ name: python3
+translation:
+ title: 可信的政府政策
+ headings:
+ Overview: 概览
+ Overview::What we build: 我们要构建什么
+ The repeated economy: 重复经济
+ Recursive strategies and promised values: 递归策略与承诺值
+ Recursive strategies and promised values::Historical antecedents: 历史渊源
+ The Abreu–Pearce–Stacchetti method: 阿布鲁-皮尔斯-斯塔凯蒂方法
+ The Abreu–Pearce–Stacchetti method::Computing the operator: 计算算子
+ The best and the worst equilibrium: 最优与最差均衡
+ The best and the worst equilibrium::The worst: 最差
+ The best and the worst equilibrium::The best: 最优
+ Examples of recursive equilibria: 递归均衡示例
+ Examples of recursive equilibria::Infinite repetition of Nash: 纳什结果的无限重复
+ Examples of recursive equilibria::Infinite repetition of something better: 更优结果的无限重复
+ 'Examples of recursive equilibria::Something worse: Abreu''s stick and carrot': 更差的均衡:阿布鲁的大棒加胡萝卜
+ Multiplicity: 多重性
+ Numerical example: 数值示例
+ Interpretations: 解读
+ Interpretations::Whose expectations are they?: 到底是谁的预期?
+ Interpretations::Remedies: 补救措施
+ Interpretations::Where this leaves us: 我们由此得出的结论
+ Exercises: 练习
+---
+
+(phillips_credible_policies)=
+```{raw} jupyter
+
+```
+
+# 可信的政府政策
+
+```{index} single: Phillips Curve; Credible Policies
+```
+
+```{contents} Contents
+:depth: 2
+```
+
+## 概览
+
+{doc}`phillips_credibility` 让政府陷入了一个陷阱。
+
+如果没有承诺自身行为的技术手段,一个每期都重新优化的政府最终会落入纳什结果,以 $\theta U^*$ 的速率通货膨胀,却一无所获。
+
+那是一个单期的故事,它招致了一个显而易见的反驳:一个预期明天将再次面对同一公众的政府,是有"声誉"需要维护的。
+
+本讲座遵循 {cite}`Sargent1999` 第4章的思路,认真对待这一反驳意见。
+
+我们让基德兰德-普雷斯科特经济永远重复,让政府当期的行动依赖于整个结果历史,并追问哪些结果路径可以作为**子博弈完美均衡**得到支持。
+
+答案并不是反驳意见所预期的那个。
+
+声誉既没有拯救拉姆齐结果,也没有确认纳什结果。
+
+它带来了均衡值的一个*连续统*,其中一些优于纳什结果,一些则差得多,而模型内部没有任何原则可以在其中做出选择。
+
+这正是本讲座旨在确立的结论,萨金特用一句话表达出来:
+
+> 我的结论是,可信计划的多重性用不可知论取代了悲观主义。
+
+这对整个系列讲座的论证都很重要。
+
+{doc}`phillips_two_stories` 反对*自然率理论胜利*这一说法的部分理由,正是建立在这一弱点之上:一个具有如此多均衡的理论几乎不能做出任何预测,因此依赖于政策制定者已经学到了正确均衡的故事,实际上是在为理论本身无法完成的工作做辩护。
+
+### 我们要构建什么
+
+所使用的机制是阿布鲁、皮尔斯和斯塔凯蒂 {cite}`Abreu,APS1990` 的递归方法,它计算的不是一个最优*值*,而是整个均衡值的*集合*。
+
+普通的动态规划迭代一个将延续*值*映射为值的算子。
+
+APS 迭代的算子将延续值的*集合*映射为值的*集合*,而子博弈完美均衡值的集合正是它的最大不动点。
+
+```{note}
+将一个承诺值作为状态变量,然后在该状态下进行动态规划,这种技巧有时被称为**双重动态规划**。
+
+这是 QuantEcon 关于 [斯塔克尔伯格计划](https://python-advanced.quantecon.org/dyn_stack.html) 讲座以及 {cite}`Ljungqvist2012` 中拉姆齐问题递归处理方法的核心思想。
+
+本讲座是一个精简的示例:状态是一个承诺值,我们所计算的对象是一个集合。
+```
+
+接下来我们以三种方式使用这一机制。
+
+我们通过猜测-验证法构造特定的均衡——纳什结果的无限重复、某种更优结果的无限重复,以及阿布鲁的"大棒加胡萝卜"策略,这一策略比纳什结果*更差*。
+
+我们通过两个小型规划问题计算最优和最差均衡值,并将其与 APS 算子的直接迭代进行核对。
+
+我们还展示了三种截然不同却都达到相同最差值的均衡,这是均衡概念约束力有多弱的最鲜明例证。
+
+```{note}
+萨金特本人给读者的建议值得转达:如果读者之前没有接触过这一理论,本章会显得困难;愿意接受其结论——即可信政策理论带来的是不可知论而非预测——的读者,可以直接跳转到 {doc}`phillips_adaptive`,而不会失去论证的线索。
+```
+
+让我们从导入模块开始:
+
+```{code-cell} ipython3
+import matplotlib.pyplot as plt
+import numpy as np
+from typing import NamedTuple
+```
+
+## 重复经济
+
+单期经济就是 {doc}`phillips_credibility` 中的那个经济。
+
+令 $(U, y, x)$ 分别为失业率、通货膨胀率和公众对通货膨胀的预期。
+
+失业率遵循预期增广菲利普斯曲线 $U = U^* - \theta(y - x)$,将政府的单期收益写成 $(x, y)$ 的函数为
+
+```{math}
+:label: cp_r
+
+r(x, y) = -\frac{1}{2}\left[\bigl(U^* - \theta(y - x)\bigr)^2 + y^2\right].
+```
+
+政府对预期 $x$ 的单期最优反应为
+
+```{math}
+:label: cp_B
+
+B(x) = \frac{\theta\,(U^* + \theta x)}{\theta^2 + 1},
+```
+
+纳什结果为 $y^N = \theta U^*$,拉姆齐结果为 $y^R = 0$。
+
+现在有两点新内容。
+
+第一,经济在 $t = 1, 2, \ldots$ 不断重复,政府按照下式对结果路径 $(x, y) = \{x_t, y_t\}_{t=1}^\infty$ 排序:
+
+```{math}
+:label: cp_value
+
+V^g(x, y) = (1 - \delta)\sum_{t=1}^{\infty}\delta^{t-1} r(x_t, y_t),
+\qquad \delta \in (0, 1).
+```
+
+因子 $(1-\delta)$ 使 $V^g$ 与单期收益处于相同的单位,从而使值与单期回报可以直接比较。
+
+第二,通货膨胀被限制在一个有界区间 $Y = [0, y^\#]$ 内。
+
+下界是拉姆齐结果;上界 $y^\#$ 使政府的问题变得非平凡,我们在寻找*最差*均衡时会看到这一点。
+
+```{code-cell} ipython3
+class Model(NamedTuple):
+ θ: float = 1.25 # slope of the Phillips curve
+ U_star: float = 5.5 # natural rate of unemployment
+ y_max: float = 10.0 # y^#, the highest admissible inflation rate
+ δ: float = 0.95 # discount factor
+
+
+def r(m, x, y):
+ "One-period government payoff when the public expects x and inflation is y."
+ U = m.U_star - m.θ * (y - x)
+ return -0.5 * (U**2 + y**2)
+
+
+def B(m, x):
+ "The government's one-period best response to expected inflation x."
+ return m.θ * (m.U_star + m.θ * x) / (m.θ**2 + 1)
+
+
+def y_nash(m):
+ return m.θ * m.U_star
+```
+
+下面所有的推导都依赖于两个收益方案,两者都是公众已形成预期的单一通货膨胀率 $y$ 的函数。
+
+第一个是**理性预期收益** $r(y, y)$:如果政府实现了公众所预期的通货膨胀,它得到的收益。
+
+第二个是**背离收益** $r(y, B(y))$:如果公众预期为 $y$,而政府屈服于诱惑做出最优反应,它得到的收益。
+
+两者都有值得记录的闭合形式,因为它们解释了后面的一切:
+
+```{math}
+:label: cp_schedules
+
+r(y, y) = -\frac{1}{2}\left(U^{*2} + y^2\right),
+\qquad
+r\bigl(y, B(y)\bigr) = -\frac{\left(U^* + \theta y\right)^2}{2\left(1 + \theta^2\right)} .
+```
+
+```{code-cell} ipython3
+def r_keep(m, y):
+ "Payoff from delivering the expected inflation rate: r(y, y)."
+ return -0.5 * (m.U_star**2 + y**2)
+
+
+def r_cheat(m, y):
+ "Payoff from best-responding to an expectation of y: r(y, B(y))."
+ return -(m.U_star + m.θ * y)**2 / (2 * (1 + m.θ**2))
+
+
+m = Model()
+ys = np.linspace(0, m.y_max, 7)
+print("closed forms agree with the definitions:",
+ np.allclose(r_keep(m, ys), r(m, ys, ys)),
+ np.allclose(r_cheat(m, ys), r(m, ys, B(m, ys))))
+```
+
+两个方案都随 $y$ 上升而下降,但下降速率不同,且恰好在纳什速率处相交,因为在该处背离的诱惑消失了,$B(y^N) = y^N$。
+
+```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "两个收益方案:实现预期通货膨胀与背离预期通货膨胀"
+ name: fig-cpol-schedules
+---
+grid = np.linspace(0, m.y_max, 400)
+yN = y_nash(m)
+v_R, v_N = r_keep(m, 0.0), r_keep(m, yN)
+v_lo = r_cheat(m, m.y_max)
+
+fig, ax = plt.subplots(figsize=(7.5, 5.5))
+ax.plot(grid, r_keep(m, grid), 'C0', lw=1.6, label='$r(y, y)$: deliver what is expected')
+ax.plot(grid, r_cheat(m, grid), 'C1', lw=1.6, label='$r(y, B(y))$: deviate')
+ax.plot(0.0, v_R, 'ko', ms=6)
+ax.annotate('$v^R$', (0.0, v_R), textcoords='offset points', xytext=(8, 4))
+ax.plot(yN, v_N, 'ko', ms=6)
+ax.annotate('$v^N$', (yN, v_N), textcoords='offset points', xytext=(8, 4))
+ax.plot(m.y_max, v_lo, 'ko', ms=6)
+ax.annotate('$v_{min}$', (m.y_max, v_lo), textcoords='offset points',
+ xytext=(-34, 4))
+ax.axvline(yN, color='k', lw=0.6, ls=':')
+ax.set_xlabel('inflation rate $y$ expected by the public')
+ax.set_ylabel('one-period payoff')
+ax.set_title('Figure 4.1: the two payoff schedules and the worst equilibrium value')
+ax.legend(loc='lower left')
+plt.show()
+```
+
+两条曲线之间的垂直差距,是政府在通货膨胀率低于预期水平时的单期诱惑大小。
+
+在 $y^N$ 处该差距为零,随着预期通货膨胀率上升超过纳什速率,差距逐渐扩大——这也是为什么*最差*均衡最终会落在 $Y$ 的顶端。
+
+## 递归策略与承诺值
+
+政府的策略必须能够使今天的行动依赖于整个过去。
+
+携带完整的历史是不可管理的,因此我们效仿 {cite}`Sargent1999`,将注意力限制在具有递归表示的策略上——这一限制不会带来任何损失,因为它不会排除任何均衡*值*。
+
+```{prf:definition} 递归政府策略
+:label: cp_strategy
+
+一个递归政府策略是一对函数 $\sigma = (\sigma_1, \sigma_2)$,连同一个初始条件 $v_1$,具有以下结构:
+
+$$
+v_1 \in \mathbb{R} \text{ given}, \qquad
+y_t = \sigma_1(v_t), \qquad
+v_{t+1} = \sigma_2(v_t, x_t, y_t),
+$$
+
+其中 $v_t$ 是一个状态变量,用于概括 $t$ 之前结果的历史。
+```
+
+私人部门的每个成员都知道 $v_t$ 及策略 $\sigma$,因此预测
+
+```{math}
+:label: cp_expect
+
+x_t = \sigma_1(v_t) .
+```
+
+方程 {eq}`cp_expect` 将理性预期内置到私人部门中:公众在均衡路径上永远不会受到意外。
+
+一个策略 $(\sigma, v_1)$ 生成一整条结果路径,从而通过 {eq}`cp_value` 生成一个值 $V^g(\sigma, v_1)$。
+
+可信政策理论此刻做了一件乍看之下像是把戏的事情。
+
+它通过要求状态变量*就是*它所生成的值,将过去与未来联系起来,
+
+```{math}
+:label: cp_fixedpt
+
+v = V^g(\sigma, v).
+```
+
+于是 $v$ 同时承担两份职责。
+
+在 $v_{t+1} = \sigma_2(v_t, x_t, y_t)$ 中,它是一个记录已发生事情的记账工具。
+
+在 {eq}`cp_fixedpt` 中,它是一个**承诺值**——政府在进入这一期时所应得的贴现未来。
+
+以第二种方式来理解,正是使这一机制运作起来的关键:策略给予政府当前和未来的结果,使其*希望*去做别人对它的预期。
+
+```{note}
+$\sigma$ 有两种解读,在一个均衡内部,二者是无法区分的。
+
+它既可以是政府所选择的决策规则,也可以是政府所遵循的公众预期体系的描述。
+
+我们会在解释一节回到这一含糊之处,因为正是从这里产生了该理论的不可知论。
+```
+
+### 历史渊源
+
+将可信性形式化是二十世纪八十年代的成就,但对这一思想的精深理解要古老得多。
+
+1784年,财政大臣雅克·内克尔向路易十六解释,为什么一位可以随意拖欠债务的绝对君主发现借款很难 {cite:p}`SargentVelde1995`:
+
+> 因此,只有通过对君主意图给予保证,并证明没有任何动机能诱使他违背其义务,才能重新点燃或维持公众信任。
+
+这句话的每一个从句都出现在现代定义中。
+
+一位君主要证明他永远不会违背自己的义务,办法是永远不*想*违背它们——通过遵从一套自带激励约束的公众预期体系,使得在每一个日期、每一种情形下,他若确认预期而非令预期落空,能获得更高的当期收益加延续值。
+
+```{prf:definition} 子博弈完美均衡
+:label: cp_spe
+
+一个具有承诺值 $v$ 的递归策略是**子博弈完美均衡**(SPE),当且仅当
+
+(a) 对于每一个 $\eta \in Y$,$\sigma_2(v, \sigma_1(v), \eta)$ 本身就是某个子博弈完美均衡可达到的值;且
+
+(b) 记 $y = \sigma_1(v)$,则
+
+$$
+v = (1-\delta)\, r(y, y) + \delta\, \sigma_2(v, y, y)
+ \;\geq\; (1-\delta)\, r(y, \eta) + \delta\, \sigma_2(v, y, \eta),
+ \qquad \forall\, \eta \in Y .
+$$
+```
+
+该定义为一个均衡附加了四个对象:一个承诺值 $v$;一个第一期结果 $(y, y)$;若遵从规定结果被观察到,一个延续值 $v'$;以及若未被观察到,另一个延续值 $\tilde v$。
+
+按照这些记号,条件 (b) 可写为
+
+$$
+v = (1-\delta) r(y, y) + \delta v' \;\geq\; (1-\delta) r(y, \eta) + \delta \tilde v,
+\qquad \forall \eta \in Y,
+$$
+
+它简单地说明政府遵从要比背离做得更好。
+
+条件 (a) 说明延续值本身必须是均衡值。
+
+该定义是循环的——均衡值同时出现在等式两侧——而这种循环性正是递归性所换来的,也是 APS 学会加以利用的东西。
+
+## 阿布鲁-皮尔斯-斯塔凯蒂方法
+
+动态规划通过迭代贝尔曼方程来计算最优值函数,这是一个将明天的值函数转化为今天的值函数的映射。
+
+APS 将这一思想应用于均衡*集合*。
+
+从一个候选延续值集合 $W \subset \mathbb{R}$ 开始。
+
+选取第一期理性预期结果 $(y, y)$ 以及来自 $W$ 的两个延续值:一个用来奖励遵从的 $w_1$,一个用来惩罚背离的 $w_2$。
+
+如果 $w_1$ 足够高,$w_2$ 足够低,这对值就支持 $y$,并带来以下值:
+
+$$
+w = (1-\delta) r(y, y) + \delta w_1 \;\geq\; (1-\delta) r(y, \eta) + \delta w_2,
+\qquad \forall \eta \in Y .
+$$
+
+```{prf:definition} 可容许性
+:label: cp_admissible
+
+若存在 $w_1, w_2 \in W$,使得对所有 $\eta \in Y$ 都有
+$w = (1-\delta) r(y,y) + \delta w_1 \geq (1-\delta) r(y,\eta) + \delta w_2$,
+则称 $(y, w)$ 关于延续值集合 $W$ 是**可容许的**。
+```
+
+令 $B(W)$ 收集所有可容许对的 $w$ 分量。
+
+这一构造内置了 {prf:ref}`cp_spe` 的条件 (b),但忽略了条件 (a),因为延续值是从一个任意集合中抽取的。
+
+```{prf:definition} 自我生成性
+:label: cp_selfgen
+
+若 $W \subseteq B(W)$,则称潜在延续值的集合 $W$ 是**自我生成的**。
+```
+
+自我生成集合中的每一个值都由取自同一集合的延续值来支持——这正是条件 (a)。
+
+APS 证明了 SPE 值的集合 $V$ 是最大的自我生成集合,$B$ 将紧集映射为紧集,$B$ 是单调的($W_2 \subseteq W_1$ 蕴含 $B(W_2) \subseteq B(W_1)$),并且从任意满足 $B(W_0) \subseteq W_0$ 的 $W_0$ 出发,迭代 $W_j = B(W_{j-1})$ 单调收敛于 $V = B(V)$。
+
+### 计算算子
+
+有两个简化使得在此处计算 $B$ 变得容易。
+
+因为政府的诱惑在其做出最优反应时最强,{prf:ref}`cp_admissible` 中的约束在 $\eta = B(y)$ 处约束最紧,所以我们可以用这一单一背离来替代"对所有 $\eta$"。
+
+而且,由于更低的惩罚会放松约束,我们总可以将 $w_2$ 设为 $W$ 中的最小元素。
+
+记 $W = [\underline w, \overline w]$,则只要 $W$ 中存在某个 $w_1$ 满足
+
+```{math}
+:label: cp_bound
+
+w_1 \;\geq\; \frac{(1-\delta)\left[r(y, B(y)) - r(y, y)\right]}{\delta} + \underline w
+\;\equiv\; \ell(y) ,
+```
+
+结果 $y$ 就是可容许的,它所生成的值遍历区间
+$\left[(1-\delta) r(y,y) + \delta \max(\underline w, \ell(y)),\;
+(1-\delta) r(y,y) + \delta \overline w\right]$。
+
+```{code-cell} ipython3
+def B_operator(m, W, n_grid=4001):
+ """
+ One application of the APS operator to an interval W = [w_lo, w_hi].
+
+ Returns the new interval, or None if no first-period outcome is admissible.
+ """
+ w_lo, w_hi = W
+ y = np.linspace(0.0, m.y_max, n_grid)
+ ℓ = (1 - m.δ) * (r_cheat(m, y) - r_keep(m, y)) / m.δ + w_lo # equation (10)
+ ok = ℓ <= w_hi # admissible outcomes
+ if not ok.any():
+ return None
+ keep = (1 - m.δ) * r_keep(m, y)
+ lows = keep[ok] + m.δ * np.maximum(w_lo, ℓ[ok])
+ highs = keep[ok] + m.δ * w_hi
+ return lows.min(), highs.max()
+
+
+def solve_aps(m, W0=None, tol=1e-12, max_iter=10_000):
+ "Iterate the APS operator to its largest fixed point."
+ W = (r_keep(m, m.y_max), 0.0) if W0 is None else W0
+ for it in range(max_iter):
+ W_new = B_operator(m, W)
+ if W_new is None:
+ return None, it
+ if max(abs(W_new[0] - W[0]), abs(W_new[1] - W[1])) < tol:
+ return W_new, it
+ W = W_new
+ return W, max_iter
+```
+
+我们从 $W_0 = [r(y^\#, y^\#),\, 0]$ 出发,这个集合足够大,能够包含每一个均衡值,然后进行迭代。
+
+```{code-cell} ipython3
+W, iters = solve_aps(m)
+print(f"converged in {iters} iterations")
+print(f"set of SPE values V = [{W[0]:.4f}, {W[1]:.4f}]")
+```
+
+## 最优与最差均衡
+
+APS 的迭代过程具有一般性,但对于*为什么*该集合最终落在这个位置,它没有告诉我们太多。
+
+对于这一经济,两个端点都可以手工求得,其论证具有启发性。
+
+### 最差
+
+最差均衡值满足以下方程:
+
+$$
+\underline v = \min_{y \in Y,\; v_1 \in V}\ \left[(1-\delta) r(y,y) + \delta v_1\right]
+\quad\text{subject to}\quad
+(1-\delta) r(y,y) + \delta v_1 \geq (1-\delta) r(y, B(y)) + \delta \underline v ,
+$$
+
+在这里,*最差*值被用作背离时的延续值——这是最严厉的惩罚。
+
+最小值在约束刚好约束(constraint binds)时取得,此时方程两侧都简化为对某个 $y$ 有 $\underline v = r(y, B(y))$。
+
+因此,找到最差均衡问题就归结为一个一维问题:
+
+```{math}
+:label: cp_worst
+
+\underline v = \min_{y \in Y}\ r\bigl(y, B(y)\bigr)
+= -\frac{\left(U^* + \theta y^\#\right)^2}{2\left(1 + \theta^2\right)} ,
+```
+
+其中闭合形式源自 {eq}`cp_schedules`,最小化行动是 $y^\#$,因为背离收益随 $y$ 单调下降。
+
+现在上界 $y^\#$ 的作用清晰起来。
+
+它是使最差惩罚保持有限的原因;若无此上界,政府可能受到任意糟糕结果的威胁,几乎任何东西都可以被支持。
+
+```{prf:proposition} 最差 SPE 是自我实施的
+:label: cp_prop_worst
+
+在最差均衡中,背离之后的延续值等于初始承诺值,因此一次背离只是简单地重新开始均衡。
+```
+
+### 最优
+
+给定 $\underline v$ 作为威胁,最优值满足
+
+```{math}
+:label: cp_best
+
+\overline v = \max_{y \in Y}\ r(y, y)
+\quad\text{subject to}\quad
+r(y, y) \geq (1-\delta) r\bigl(y, B(y)\bigr) + \delta \underline v ,
+```
+
+这里我们利用了这样一个事实:最优值必须以自身来奖励遵从,因此
+$\overline v = (1-\delta) r(y,y) + \delta \overline v = r(y,y)$。
+
+```{prf:proposition} 最优 SPE 是自我奖励的
+:label: cp_prop_best
+
+在最优均衡中,遵从之后的延续值等于承诺值。
+```
+
+```{code-cell} ipython3
+def worst_value(m):
+ "The worst SPE value and the action that attains it."
+ return r_cheat(m, m.y_max), m.y_max
+
+
+def best_value(m, v_lo, n_grid=200_001):
+ "The best SPE value, given the worst value as the punishment threat."
+ y = np.linspace(0.0, m.y_max, n_grid)
+ feasible = r_keep(m, y) >= (1 - m.δ) * r_cheat(m, y) + m.δ * v_lo
+ if not feasible.any():
+ return None, None
+ vals = np.where(feasible, r_keep(m, y), -np.inf)
+ k = vals.argmax()
+ return vals[k], y[k]
+
+
+v_lo, y_sharp = worst_value(m)
+v_hi, y_best = best_value(m, v_lo)
+
+print(f"worst value v_lo = {v_lo:.4f} attained with y = {y_sharp:.2f}")
+print(f"best value v_hi = {v_hi:.4f} attained with y = {y_best:.2f}")
+print(f"APS iteration gave [{W[0]:.4f}, {W[1]:.4f}]")
+print(f"the two agree: {np.allclose(W, (v_lo, v_hi), atol=1e-6)}")
+```
+
+规划问题和集合迭代结果一致,在此贴现因子下,最优均衡正是拉姆齐结果本身。
+
+## 递归均衡示例
+
+现在我们按照文献发现它们的顺序,通过猜测-验证法构造具体的均衡。
+
+### 纳什结果的无限重复
+
+最简单的均衡就是永远重复单期纳什结果。
+
+取 $v_1 = v^N = r(y^N, y^N)$,对每个 $v$ 都有 $\sigma_1(v) = y^N$,且对每个 $(v, x, y)$ 都有
+$\sigma_2(v, x, y) = v^N$。
+
+条件 (a) 由构造自动满足,而条件 (b) 简化为
+$r(y^N, y^N) \geq r(y^N, B(y^N))$,此式*恰好以等式成立*,因为 $y^N$ 是最优反应映射的不动点。
+
+这里没有任何东西能约束政府,也不需要:它已经在做自己最想做的事了。
+
+### 更优结果的无限重复
+
+设 $v^b = r(y^b, y^b) > v^N$ 为一个优于纳什的值,假设
+
+```{math}
+:label: cp_bg
+
+r\bigl(y^b, B(y^b)\bigr) - r(y^b, y^b)
+\;\leq\; \frac{\delta}{1-\delta}\left(v^b - v^N\right).
+```
+
+左边是背离的单期收益;右边是永久回归纳什带来的贴现损失。
+
+当 {eq}`cp_bg` 成立时,规定 $y^b$(当承诺值为 $v^b$ 时)并在任何背离后回归纳什的策略是一个 SPE。
+
+{cite:t}`BarroGordon1983` 研究了 $y^b = y^R$ 的情形:预期中的回归纳什永远支持拉姆齐结果。
+
+至于它是否成立取决于耐心程度,对于这个经济体,其阈值具有一个非常简洁的形式。
+
+在 {eq}`cp_bg` 中设 $y^b = 0$,并使用闭合形式 {eq}`cp_schedules`,每一个 $U^*$ 和每一个尺度因子都会消去,只剩下
+
+```{math}
+:label: cp_cutoff
+
+\delta \;\geq\; \delta^\star = \frac{1}{2 + \theta^2} .
+```
+
+```{code-cell} ipython3
+def supports_ramsey_by_nash(m):
+ "Does reversion to Nash sustain Ramsey forever?"
+ gain = r_cheat(m, 0.0) - r_keep(m, 0.0)
+ loss = m.δ / (1 - m.δ) * (r_keep(m, 0.0) - r_keep(m, y_nash(m)))
+ return gain <= loss
+
+
+δ_star = 1 / (2 + m.θ**2)
+print(f"closed-form cutoff δ* = 1/(2 + θ²) = {δ_star:.6f}\n")
+for δ in (0.20, 0.27, δ_star - 1e-6, δ_star + 1e-6, 0.50, 0.95):
+ md = m._replace(δ=δ)
+ print(f" δ = {δ:.8f}: Nash reversion supports Ramsey? "
+ f"{supports_ramsey_by_nash(md)}")
+```
+
+在 $\delta^\star \approx 0.28$ 以上,回归纳什的威胁足以将政府维持在零通货膨胀水平;在此之下,威胁则过于薄弱。
+
+### 更差的均衡:阿布鲁的大棒加胡萝卜
+
+当回归纳什的威胁不够强时,{cite:t}`Abreu` 探讨了是否可以用某个*更差*的均衡作为惩罚。
+
+他的做法是一种大棒加胡萝卜策略:要求政府以最大速率 $y^\#$ 通货膨胀一期——这是大棒——之后奖励它以永远的拉姆齐结果。
+
+此值为
+
+```{math}
+:label: cp_abreu
+
+\tilde v = (1-\delta)\, r(y^\#, y^\#) + \delta\, v^R ,
+```
+
+拒绝接受大棒的惩罚是*重新开始*整个策略,因此背离时的延续值就是 $\tilde v$ 本身。
+
+正是这种自我指涉的选择,使得计算变得如此简单。
+
+将 $\tilde v$ 代入激励约束的两侧,$\delta$ 项会相消,该策略恰好在
+$\tilde v \geq r\bigl(y^\#, B(y^\#)\bigr) = \underline v$ 时是一个均衡。
+
+```{code-cell} ipython3
+def abreu_value(m):
+ "Value of the stick-and-carrot strategy: one period at y^#, then Ramsey."
+ return (1 - m.δ) * r_keep(m, m.y_max) + m.δ * r_keep(m, 0.0)
+
+
+for δ in (0.20, 0.95):
+ md = m._replace(δ=δ)
+ va, vn = abreu_value(md), r_keep(md, y_nash(md))
+ print(f"δ = {δ}: v_abreu = {va:8.4f} v^N = {vn:8.4f} "
+ f"worse than Nash? {va < vn} is an SPE? {va >= worst_value(md)[0]}")
+```
+
+在 $\delta = 0.95$ 时,大棒加胡萝卜策略的值反而*优于*纳什,因为单个糟糕的时期被永久的拉姆齐结果远远抵消了。
+
+在 $\delta = 0.2$ 时,它则远*差于*纳什——这正是要点所在。
+
+比纳什更差的惩罚是比回归纳什更强的威慑,因此它可以支持纳什回归所无法支持的结果。
+
+```{code-cell} ipython3
+m2 = m._replace(δ=0.20)
+
+def supports_ramsey_with_threat(m, threat):
+ "Does the threat value sustain Ramsey forever?"
+ return r_keep(m, 0.0) >= (1 - m.δ) * r_cheat(m, 0.0) + m.δ * threat
+
+print(f"at δ = 0.2:")
+print(f" threat = Nash ({r_keep(m2, y_nash(m2)):8.4f}): "
+ f"supports Ramsey? {supports_ramsey_with_threat(m2, r_keep(m2, y_nash(m2)))}")
+print(f" threat = Abreu ({abreu_value(m2):8.4f}): "
+ f"supports Ramsey? {supports_ramsey_with_threat(m2, abreu_value(m2))}")
+```
+
+在 $\delta = 0.2$ 时,一个缺乏耐心的政府无法被纳什的前景约束住,但*可以*被阿布鲁大棒的前景约束住。
+
+有两个定义可以命名我们刚刚遇到的结构。
+
+```{prf:definition} 自我实施与自我奖励
+:label: cp_selfenf
+
+一个递归 SPE 若背离后的延续值等于初始承诺值,则称为**自我实施的**;若遵从后的延续值等于承诺值,则称为**自我奖励的**。
+```
+
+阿布鲁的大棒加胡萝卜策略是自我实施的;对优于纳什的结果进行无限重复则是自我奖励的。
+
+{prf:ref}`cp_prop_worst` 和 {prf:ref}`cp_prop_best` 表明,这两种结构并非奇特案例:它们恰恰就是极端均衡的样子。
+
+## 多重性
+
+这一理论内蕴含着两层多重性。
+
+存在一个均衡*值*的连续统,而且——更为尖锐的一点是——许多不同的结果*路径*都能达到同一个值。
+
+为了说明第二点,我们构造了三个均衡,它们都得到最差值 $\underline v$。
+
+三者都以相同的方式开始:第一期承诺值为 $\underline v$,规定的行动是 $y^\#$,延续值 $v_2$ 满足
+$\underline v = (1-\delta) r(y^\#, y^\#) + \delta v_2$。
+
+它们的区别在于此后的做法。
+
+```{code-cell} ipython3
+def next_value(m, v):
+ "Solve v = (1-δ) r(y^#, y^#) + δ v' for the continuation value v'."
+ return (v - (1 - m.δ) * r_keep(m, m.y_max)) / m.δ
+
+
+def action_delivering(m, v):
+ "The inflation rate ỹ with r(ỹ, ỹ) = v."
+ return np.sqrt(max(-2.0 * v - m.U_star**2, 0.0))
+
+
+def method_1(m, v_lo, v_hi, max_steps=500):
+ "Climb at y^# until one more step would overshoot v_hi, then settle at v_hi."
+ v, y = [v_lo], []
+ for _ in range(max_steps):
+ nxt = next_value(m, v[-1])
+ if nxt > v_hi:
+ break
+ y.append(m.y_max)
+ v.append(nxt)
+ target = (v[-1] - m.δ * v_hi) / (1 - m.δ) # r(ỹ, ỹ) for the switching period
+ y.append(action_delivering(m, target))
+ v.append(v_hi)
+ y.append(action_delivering(m, v_hi)) # stay at v_hi forever after
+ return np.array(v), np.array(y)
+
+
+def method_2(m, v_lo, max_steps=500):
+ "Climb at y^# only until the promised value first exceeds v^N, then freeze."
+ v_N = r_keep(m, y_nash(m))
+ v, y = [v_lo], []
+ for _ in range(max_steps):
+ nxt = next_value(m, v[-1])
+ y.append(m.y_max)
+ v.append(nxt)
+ if nxt > v_N:
+ break
+ y.append(action_delivering(m, v[-1])) # freeze at v** forever
+ return np.array(v), np.array(y)
+
+
+def method_3(m, v_lo):
+ "One period at y^#, then freeze immediately at v_2."
+ v_2 = next_value(m, v_lo)
+ return (np.array([v_lo, v_2, v_2]),
+ np.array([m.y_max, action_delivering(m, v_2),
+ action_delivering(m, v_2)]))
+```
+
+```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 达到相同最差值的三种子博弈完美均衡
+ name: fig-cpol-three-methods
+---
+paths = {'method 1': method_1(m, v_lo, v_hi),
+ 'method 2': method_2(m, v_lo),
+ 'method 3': method_3(m, v_lo)}
+
+fig, axes = plt.subplots(2, 3, figsize=(13, 6.5), sharex='col')
+for k, (name, (v, y)) in enumerate(paths.items()):
+ axes[0, k].plot(range(1, len(v) + 1), v, 'C0o-', ms=3.5, lw=1)
+ axes[0, k].axhline(v_lo, color='k', ls='--', lw=0.8)
+ axes[0, k].axhline(v_hi, color='C2', ls=':', lw=0.8)
+ axes[0, k].set_title(f'{name}: continuation values', fontsize=10)
+ axes[1, k].plot(range(1, len(y) + 1), y, 'C1o-', ms=3.5, lw=1)
+ axes[1, k].set_ylim(-0.4, m.y_max + 0.4)
+ axes[1, k].set_xlabel('$t$')
+ axes[1, k].set_title(f'{name}: inflation', fontsize=10)
+axes[0, 0].set_ylabel('promised value $v_t$')
+axes[1, 0].set_ylabel('inflation $y_t$')
+fig.suptitle('Figures 4.2-4.4: three equilibria that all attain the worst '
+ 'value $v_{min}$')
+plt.tight_layout()
+plt.show()
+```
+
+```{code-cell} ipython3
+print(f"Nash value {r_keep(m, y_nash(m)):.4f} at inflation {y_nash(m):.4f}; "
+ f"worst value {v_lo:.4f}\n")
+for name, (v, y) in paths.items():
+ print(f"{name:9s}: {len(y):3d} periods before settling, "
+ f"first-period value {v[0]:.4f}, "
+ f"terminal value {v[-1]:8.4f} at inflation {y[-1]:.4f}")
+```
+
+这三条路径几乎完全不同。
+
+第一条路径以最大通货膨胀速率攀升约六十期,然后才降至拉姆齐结果。
+
+第二条路径攀升的时间较短,随后永远冻结在一个刚刚*优于*纳什的值上,由一个略低于纳什速率的通货膨胀率所支撑。
+
+第三条路径以最大速率通货膨胀一期,然后立即稳定下来,其值只比最差值高一点点。
+
+三者都是子博弈完美的,且三者给政府带来的值完全相同。
+
+均衡概念对于这三者中哪一个描述了真实世界,无话可说。
+
+## 数值示例
+
+以下是 {cite}`Sargent1999` 第4章报告的数值示例,我们从头进行了计算。
+
+```{code-cell} ipython3
+import pandas as pd
+
+rows = [
+ ("$\\theta$", m.θ, ""),
+ ("$U^*$", m.U_star, ""),
+ ("$y^\\#$", m.y_max, ""),
+ ("$\\delta$", m.δ, ""),
+ ("$y^N$ (Nash inflation)", y_nash(m), "6.8750"),
+ ("$y^R$ (Ramsey inflation)", 0.0, "0"),
+ ("$v^R$", r_keep(m, 0.0), "-15.1250"),
+ ("$v^N$", r_keep(m, y_nash(m)), "-38.7578"),
+ ("$\\underline v$", v_lo, "-63.2195"),
+ ("$v_{\\rm abreu}$", abreu_value(m), "-17.6250"),
+]
+pd.DataFrame(rows, columns=["object", "computed", "reported in the book"]).set_index("object")
+```
+
+```{code-cell} ipython3
+print(f"cutoff discount factor δ* : {1 / (2 + m.θ**2):.4f} (book reports 0.2807)")
+print(f"Abreu stick-and-carrot at δ = 0.2: {abreu_value(m._replace(δ=0.2)):.4f}"
+ f" (book reports -55.125)")
+```
+
+每一项都与书中一致。
+
+## 解读
+
+关于可信计划的文献,对于 {doc}`phillips_two_stories` 中*自然率理论胜利*的说法喜忧参半。
+
+它确实将政府从 {doc}`phillips_credibility` 的悲观境地中拯救了出来:优于纳什的结果是可以实现的,而且在合理的贴现因子下,拉姆齐结果本身就是一个均衡。
+
+但它拯救得太多了。
+
+比纳什更差的值也是均衡,而两个极端之间还存在一个连续统。
+
+众多结果在经验上使模型变得不明确,也消解了早期理性预期研究者从这一假设中所期望得到的东西——消除描述预期的自由参数。
+
+### 到底是谁的预期?
+
+在1979年对保罗·麦克拉肯编辑的一份经合组织(OECD)报告的评论中,卢卡斯抗议该报告的建议——"各国政府应努力促进良好的预期",仿佛预期是一套额外的政策工具。
+
+在1979年,这一抗议是有的放矢的:当时理性预期模型将政府政策视为外生的,并将预期作为政府政策的一个*函数*,通过跨方程约束加以关联。
+
+可信政策理论以某种方式改变了这一局面,但对双方都不能算作清楚的支持。
+
+它将预期体系变成了会影响结果的自由参数——但这些参数是模型内部没有任何人可以选择的。
+
+政府遵从关于其自身行为的均衡预期。
+
+在一个均衡内部,政府的策略同时是一条决策规则,也是对公众预期的描述,而两者是无法分开的。
+
+正如萨金特所言:麦克拉肯报告的作者们相信多重性和操纵性,而卢卡斯对两者都持怀疑态度;关于可信计划的文献支持多重性,但不支持操纵性。
+
+### 补救措施
+
+单凭声誉是反通货膨胀政策的一个薄弱基础,这种薄弱之处激发了改变博弈规则、而非寄望于博弈内部出现一个好均衡的提议。
+
+{cite:t}`Rogoff1985` 提议将货币政策委托给一个比社会整体更不关心失业的人。
+
+将其委托给一个甚至不了解*暂时性*权衡取舍的人,同样有效;艾伦·布林德后来提出了一个相关的方案:委托给一个了解自然率、且从不希望失业率偏离该水平的当局。
+
+维持一个具有不同通货膨胀-失业偏好的潜在中央银行家人才库,也能改善结果 {cite:p}`BarroGordon1983`。
+
+### 我们由此得出的结论
+
+上述每一项补救措施改变的都是*制度*,而非理论本身。
+
+在理论内部,多重性是不可消解的,萨金特精准地指出了它的来源:它源于赋予系统中*每一个人*的理性。
+
+之所以会产生这一连续统,正是因为完美理性——一套足以支持某一均衡的复杂预期体系,也足以支持许多均衡。
+
+因此本书从完美理性中撤退。
+
+后续的讲座将完全理性的参与者替换为对经济理解有限的参与者——首先是 {doc}`phillips_adaptive` 中的公众,然后是 {doc}`phillips_misspecified` 和 {doc}`phillips_self_confirming` 中的双方,最后是 {doc}`phillips_learning` 中实时估计其模型的政府。
+
+这些模型更接近卢卡斯批评麦克拉肯报告时心目中所想的东西,而且与本讲座的理论不同,它们能够做出预测。
+
+## 练习
+
+```{exercise-start}
+:label: cpol_ex1
+```
+
+阈值 {eq}`cp_cutoff` 声称:当且仅当 $\delta \geq 1/(2 + \theta^2)$ 时,回归纳什能永远维持拉姆齐结果——这一阈值只依赖于菲利普斯曲线的斜率,而不依赖于自然率 $U^*$ 或上界 $y^\#$。
+
+请通过数值方法验证这两个说法。
+
+对于一系列 $\theta$ 值,在细网格上找出能使回归纳什支持拉姆齐结果的最小 $\delta$,并与 $1/(2+\theta^2)$ 进行比较。
+
+然后检验改变 $U^*$ 是否会影响这一结果。
+
+```{exercise-end}
+```
+
+```{solution-start} cpol_ex1
+:class: dropdown
+```
+
+```{code-cell} ipython3
+δ_grid = np.linspace(0.01, 0.99, 9801)
+
+rows = []
+for θ in (0.5, 1.0, 1.25, 2.0, 3.0):
+ for U_star in (5.5, 11.0):
+ md = Model(θ=θ, U_star=U_star)
+ ok = [δ for δ in δ_grid
+ if supports_ramsey_by_nash(md._replace(δ=δ))]
+ rows.append((θ, U_star, min(ok), 1 / (2 + θ**2)))
+
+pd.DataFrame(rows, columns=["$\\theta$", "$U^*$", "smallest $\\delta$ found",
+ "$1/(2+\\theta^2)$"]).round(4)
+```
+
+网格搜索的结果与闭合形式的匹配误差不超过网格间距,且 $U^*$ 加倍不改变任何结果。
+
+$U^*$ 和 $y^\#$ 都被消去的原因,在 {eq}`cp_schedules` 中是可见的:$y = 0$ 处的单期诱惑以及纳什-拉姆齐值差距都与 $U^{*2}$ 成正比,因此尺度从不等式中消去,而 $y^\#$ 从不出现,因为不等式两侧都不涉及上界。
+
+菲利普斯曲线越陡峭——$\theta$ 越大——纳什结果相对于拉姆齐结果就越差,这会强化威胁,使得一个*更缺乏耐心*的政府也能被维持在零通货膨胀水平。
+
+```{solution-end}
+```
+
+```{exercise-start}
+:label: cpol_ex2
+```
+
+随着政府变得缺乏耐心,均衡值的集合会收缩。
+
+使用 `solve_aps` 计算一系列贴现因子下的均衡值集合 $V = [\underline v, \overline v]$,并将两个端点相对 $\delta$ 作图,同时标出纳什值和拉姆齐值。
+
+大致在什么样的贴现因子下,最优均衡值不再是拉姆齐结果?
+
+当 $\delta \to 0$ 时,这一集合会发生什么变化?为什么这正是你应该预料到的答案?
+
+```{exercise-end}
+```
+
+```{solution-start} cpol_ex2
+:class: dropdown
+```
+
+```{code-cell} ipython3
+δs = np.linspace(0.02, 0.98, 49)
+lows, highs = [], []
+for δ in δs:
+ Wδ, _ = solve_aps(m._replace(δ=δ))
+ lows.append(Wδ[0])
+ highs.append(Wδ[1])
+
+fig, ax = plt.subplots(figsize=(8.5, 5))
+ax.fill_between(δs, lows, highs, alpha=0.2, color='C0', label='set of SPE values $V$')
+ax.plot(δs, highs, 'C0', lw=1.4)
+ax.plot(δs, lows, 'C0', lw=1.4)
+ax.axhline(r_keep(m, 0.0), color='C2', ls=':', lw=1.2, label='Ramsey $v^R$')
+ax.axhline(r_keep(m, y_nash(m)), color='k', ls='--', lw=1, label='Nash $v^N$')
+ax.set_xlabel(r'discount factor $\delta$')
+ax.set_ylabel('value')
+ax.set_title('SPE values by discount factor')
+ax.legend()
+plt.show()
+```
+
+```{code-cell} ipython3
+best_is_ramsey = [δ for δ, h in zip(δs, highs)
+ if np.isclose(h, r_keep(m, 0.0), atol=1e-6)]
+print(f"best value equals Ramsey for δ ≥ {min(best_is_ramsey):.3f}")
+print(f"closed-form cutoff for Nash reversion: δ* = {1/(2+m.θ**2):.3f}")
+```
+
+最优均衡值等于拉姆齐结果的下限,比 {eq}`cp_cutoff` 中约 $\delta^\star \approx 0.28$ 的贴现因子低得多,原因在于阿布鲁策略。
+
+阈值 $\delta^\star$ 是使用回归*纳什*作为威胁推导出来的;而 APS 集合使用最差均衡作为威胁,这要严厉得多,因此拉姆齐结果在比巴罗和戈登的论证单独所暗示的更低的贴现因子下依然存活。
+
+当 $\delta \to 0$ 时,该集合会坍缩至单点 $v^N$。
+
+这正是它必然要发生的:一个完全没有耐心的政府只关心当期,对未来的承诺没有任何分量,唯一能被维持的结果就是政府会短视地选择的那个结果——纳什结果。
+
+```{solution-end}
+```
+
+```{exercise-start}
+:label: cpol_ex3
+```
+
+多重性一节中的方法3是三者中最为大胆的:它以 $y^\#$ 通货膨胀一期,然后永远冻结在一个固定的通货膨胀率上。
+
+因为它立即稳定下来,其激励约束是三者中最紧的,因此值得核实一下,而不是想当然地认为它成立。
+
+请直接验证方法3是一个子博弈完美均衡:确认其第一期的承诺得到兑现,其冻结的延续值落在 $V$ 内,并确认政府在*两个*阶段都更偏好遵从而非背离,使用 $\underline v$ 作为惩罚。
+
+```{exercise-end}
+```
+
+```{solution-start} cpol_ex3
+:class: dropdown
+```
+
+```{code-cell} ipython3
+v_2 = next_value(m, v_lo)
+y_tilde = action_delivering(m, v_2)
+
+# phase 1: promised v_lo, prescribed action y^#, continuation v_2 on adherence
+lhs_1 = (1 - m.δ) * r_keep(m, m.y_max) + m.δ * v_2
+rhs_1 = (1 - m.δ) * r_cheat(m, m.y_max) + m.δ * v_lo
+
+# phase 2: promised v_2, prescribed action ỹ, continuation v_2 on adherence
+lhs_2 = (1 - m.δ) * r_keep(m, y_tilde) + m.δ * v_2
+rhs_2 = (1 - m.δ) * r_cheat(m, y_tilde) + m.δ * v_lo
+
+print(f"v_2 = {v_2:.4f}, ỹ = {y_tilde:.4f}")
+print(f"v_2 lies inside V = [{v_lo:.4f}, {v_hi:.4f}]: "
+ f"{v_lo <= v_2 <= v_hi}")
+print(f"phase 1 delivers the promise: {np.isclose(lhs_1, v_lo)}")
+print(f"phase 1 incentive: {lhs_1:.4f} >= {rhs_1:.4f} -> {lhs_1 >= rhs_1}")
+print(f"phase 2 delivers the promise: {np.isclose(lhs_2, v_2)}")
+print(f"phase 2 incentive: {lhs_2:.4f} >= {rhs_2:.4f} -> {lhs_2 >= rhs_2}")
+print(f"slack in phase 2: {lhs_2 - rhs_2:.6f}")
+```
+
+所有条件都成立,且第一期的承诺被精确地兑现。
+
+第二阶段的余量非常小,这并非偶然。
+
+冻结值 $v_2$ 只是勉强高于 $\underline v$,因此背离的惩罚只比均衡本身稍差一点——这恰恰意味着它接近均衡值集合的底部。
+
+如果进一步压低这一构造,激励约束就会失效,这也是理解 $\underline v$ 为何恰好落在此处的另一种方式。
+
+```{solution-end}
+```
\ No newline at end of file
diff --git a/lectures/phillips_drifts_volatilities.md b/lectures/phillips_drifts_volatilities.md
index d49fd21..b99ca76 100644
--- a/lectures/phillips_drifts_volatilities.md
+++ b/lectures/phillips_drifts_volatilities.md
@@ -56,6 +56,9 @@ translation:
# 漂移与波动率
+```{index} single: Phillips Curve; Drifts and Volatilities
+```
+
```{contents} Contents
:depth: 2
```
@@ -90,10 +93,9 @@ translation:
我们将依次讲解数据变换、先验、抽样器和主要的实证结果。
-让我们从一些导入语句和数据路径开始。
+让我们从一些导入语句和数据的 URL 开始。
```{code-cell} ipython3
-from pathlib import Path
import time
import matplotlib.pyplot as plt
@@ -105,19 +107,11 @@ from scipy.special import expit
from scipy.stats import invwishart
-def locate_data_assets():
- """Find assets from either a MyST build or the repository root."""
- relative = Path('_static/lecture_specific/phillips_drifts_volatilities')
- candidates = (relative, Path('lectures') / relative)
- for candidate in candidates:
- if (candidate / 'NEWQDATA.csv').is_file():
- return candidate
- searched = ', '.join(str(path.resolve()) for path in candidates)
- raise FileNotFoundError(f'NEWQDATA.csv was not found; searched {searched}')
-
-
-asset_path = locate_data_assets()
-data_path = asset_path / 'NEWQDATA.csv'
+data_url = (
+ 'https://raw.githubusercontent.com/QuantEcon/lecture-python.myst/'
+ 'main/lectures/_static/lecture_specific/phillips_drifts_volatilities/'
+ 'NEWQDATA.csv'
+)
```
## 政策不当还是运气不好?
@@ -146,12 +140,13 @@ data_path = asset_path / 'NEWQDATA.csv'
因此,科格利和萨金特构建了一个能同时容纳*这两种*渠道的模型,并让贝叶斯后验来判定数据究竟需要多少这两种成分。
+(csdv-model)=
## 一个系数漂移且波动率随机变化的 VAR
设变量按名义利率、变换后的失业率、通货膨胀的顺序排列,
$$
-y_t = \begin{bmatrix} i_t & u_t & \pi_t \end{bmatrix}'.
+y_t = \begin{bmatrix} i_t & u_t & \pi_t \end{bmatrix}^\top.
$$
(这里的 $u_t$ 并非原始的失业率,而是其logit变换,我们将在下面的数据部分定义这一变换。)
@@ -160,9 +155,9 @@ $$
```{math}
:label: csdv_measurement
-y_t = X_t'\theta_t + \varepsilon_t,
+y_t = X_t^\top \theta_t + \varepsilon_t,
\qquad
-X_t' = I_3 \otimes \begin{bmatrix} 1 & y_{t-1}' & y_{t-2}' \end{bmatrix}.
+X_t^\top = I_3 \otimes \begin{bmatrix} 1 & y_{t-1}^\top & y_{t-2}^\top \end{bmatrix}.
```
每个方程都有一个截距项和六个滞后系数,因此 $\theta_t$ 包含
@@ -324,7 +319,7 @@ $\beta$ 和 $(\sigma_1,\sigma_2,\sigma_3)$。
```{code-cell} ipython3
def prepare_data(source, ordering=('i', 'u', 'pi')):
"""Transform a quarterly table and construct the VAR data."""
- if isinstance(source, (str, Path)):
+ if isinstance(source, str):
table = pd.read_csv(source)
else:
table = source.copy()
@@ -359,7 +354,7 @@ def prepare_data(source, ordering=('i', 'u', 'pi')):
}
-data = prepare_data(data_path)
+data = prepare_data(data_url)
data_summary = pd.Series(
{
@@ -521,6 +516,7 @@ prior_summary = pd.Series(
prior_summary.to_frame()
```
+(csdv-sampler)=
## 一个 Metropolis-within-Gibbs 抽样器
我们通过遍历 {cite:t}`CogleySargent2005` 所使用的五个参数模块来模拟后验。
@@ -824,7 +820,7 @@ $$
$$
\pi_{\mathcal A}(z\mid\lambda,Y^T)
= \frac{N(z;m,C)\,\mathbb{1}_{\mathcal A}(z)}
- {\Pr(z\in\mathcal A\mid\lambda,Y^T)}.
+ {\mathbb{P}\{z\in\mathcal A\mid\lambda,Y^T\}}.
$$
那个归一化概率难以计算,但椭圆转移从不需要对它求值。
@@ -1073,10 +1069,11 @@ def validate_posterior_arrays(result, periods):
expected_shapes = validate_posterior_arrays(posterior, len(data['dates']))
```
+(csdv-results)=
## 数据揭示了什么
-我们用后验均值系数路径 $E(\theta_t\mid T)$ 和后验均值协方差路径
-$E(R_t\mid T)$ 来总结后验,然后在我们所提出问题的背景下对其加以解释。
+我们用后验均值系数路径 $\mathbb{E}(\theta_t\mid T)$ 和后验均值协方差路径
+$\mathbb{E}(R_t\mid T)$ 来总结后验,然后在我们所提出问题的背景下对其加以解释。
### 漂移的速率与结构
@@ -1465,11 +1462,11 @@ f_{\pi\pi}(\omega,t)
s_\pi
(I-A_{t\mid T}e^{-i\omega})^{-1}
\mathcal R_t
-(I-A_{t\mid T}'e^{i\omega})^{-1}
-s_\pi',
+(I-A_{t\mid T}^\top e^{i\omega})^{-1}
+s_\pi^\top,
```
-其中 $\mathcal R_t$ 将 $E(R_t\mid T)$ 嵌入伴随系统中。
+其中 $\mathcal R_t$ 将 $\mathbb{E}(R_t\mid T)$ 嵌入伴随系统中。
低频功率既取决于自回归系数,也取决于创新协方差。
@@ -1742,8 +1739,8 @@ plt.show()
```{math}
:label: csdv_policy_rule
i_t = \beta_0
-+ \beta_1 E_t\bar\pi_{t,t+h_\pi}
-+ \beta_2 E_t\bar u_{t,t+h_u}
++ \beta_1 \mathbb{E}_t\bar\pi_{t,t+h_\pi}
++ \beta_2 \mathbb{E}_t\bar u_{t,t+h_u}
+ \beta_3 i_{t-1}
+ \nu_t.
```
@@ -2076,7 +2073,7 @@ if len(incomplete_quarters):
else:
complete_extension = quarterly_unfilled
-cs_sample = pd.read_csv(data_path)
+cs_sample = pd.read_csv(data_url)
overlap_date = pd.Timestamp('2000-10-01')
cs_sample_overlap = cs_sample.iloc[-1]
latest_overlap = latest_quarterly.loc[overlap_date]
@@ -2964,6 +2961,7 @@ plt.show()
2025年第三季度的自然失业率和政策边际估计仍不够精确,尤其是因为并非每一次后验抽样都满足 $|\beta_3|<1$。
+(csdv-verdict)=
## 政策不当还是运气不好?一个结论
贝叶斯 VAR 模型对本讲座开篇提出的问题给出了一个细致入微的答案。
@@ -3087,4 +3085,4 @@ plt.show()
因此,仅靠漂移的波动率能够解释一部分上升过程,但完全不能解释沃尔克时期持续性的反通胀式下降——1980年之后的这次崩溃,正是关于*系统性*动态的证据,这也正是为何要回答政策不当还是运气不好这个问题,需要同时考虑这两个渠道。
```{solution-end}
-```
\ No newline at end of file
+```
diff --git a/lectures/phillips_escaping_nash.md b/lectures/phillips_escaping_nash.md
index 38e1d2a..fde80fe 100644
--- a/lectures/phillips_escaping_nash.md
+++ b/lectures/phillips_escaping_nash.md
@@ -37,6 +37,9 @@ translation:
# 逃离纳什通胀
+```{index} single: Phillips Curve; Escaping Nash Inflation
+```
+
```{contents} Contents
:depth: 2
```
@@ -88,7 +91,7 @@ U_n = u - \theta(\pi_n - \hat x_n) + \sigma_1 W_{1n},
\hat x_n = x_n,
```
-其中 $\theta, u > 0$,且 $W_n = (W_{1n}, W_{2n})'$ 为独立同分布的标准高斯变量。
+其中 $\theta, u > 0$,且 $W_n = (W_{1n}, W_{2n})^\top$ 为独立同分布的标准高斯变量。
政府并不知道 {eq}`en_truth`。
@@ -102,6 +105,12 @@ U_n = \gamma_1 \pi_n + \gamma_{-1} + \eta_n ,
信念为 $\gamma = (\gamma_1, \gamma_{-1})$(斜率与截距),并将 $\eta_n$ 视为外生的。
+```{note}
+这里斜率在前,回归量为 $\Phi = (\pi, 1)$,与 {cite}`ChoWilliamsSargent2002` 以及 {doc}`phillips_self_confirming` 中 $\gamma_1, \gamma_{-1}$ 的约定保持一致。
+
+{doc}`phillips_priors` 研究了同一模型,但截距在前;有关这种转换请参阅该处的提示。
+```
+
在相信 {eq}`en_belief` 的前提下,政府求解 {doc}`菲尔普斯问题 `,其静态最优反应将通货膨胀设定为常数
```{math}
@@ -146,7 +155,7 @@ class EscapeModel:
将这些总体系数记为 $T(\gamma)$,CWS 证明
$$
-\bar g(\gamma) \equiv E\left[\Phi(U - \Phi'\gamma)\right] = \bar M \left(T(\gamma) - \gamma\right),
+\bar g(\gamma) \equiv \mathbb{E}\left[\Phi(U - \Phi^\top \gamma)\right] = \bar M \left(T(\gamma) - \gamma\right),
$$
因此自我确认均衡解出 $\bar g(\gamma) = 0$。
@@ -165,9 +174,15 @@ print(f"check g_bar = {model.g_bar(γ_sce)}")
```
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: The unique self-confirming equilibrium in belief space
+ name: fig-esc-sce
+---
γ1_grid = np.linspace(-2, 1, 200)
fig, ax = plt.subplots(figsize=(7, 6))
-ax.plot(γ1_grid, model.u * (1 + γ1_grid**2), label=r'$\gamma_{-1} = u(1+\gamma_1^2)$')
+ax.plot(γ1_grid, model.u * (1 + γ1_grid**2), label=r'$\gamma_{-1} = u(1+\gamma_1^2)$', lw=2)
ax.axvline(-model.θ, color='C1', ls='--', label=r'$\gamma_1 = -\theta$')
ax.plot(γ_sce[0], γ_sce[1], 'ko', ms=8)
ax.annotate('SCE (Nash)', γ_sce, (γ_sce[0] + 0.1, γ_sce[1] + 1))
@@ -198,6 +213,12 @@ plt.show()
因此,仅在均值动态作用下,适应性政府会被吸引到纳什通货膨胀水平。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "均值动态:从一个受扰动的起点出发,信念回归纳什水平"
+ name: fig-esc-mean-dynamics
+---
def mean_ode(t, z, model):
γ, R = z[:2], z[2:].reshape(2, 2)
return np.concatenate([np.linalg.inv(R) @ model.g_bar(γ),
@@ -236,7 +257,7 @@ plt.show()
```{math}
:label: en_control
-\bar S = \inf_{v(\cdot),\, T} \; \frac12 \int_0^T v(s)' Q(\gamma(s), R(s))^{-1} v(s)\, ds
+\bar S = \inf_{v(\cdot),\, T} \; \frac12 \int_0^T v(s)^\top Q(\gamma(s), R(s))^{-1} v(s)\, ds
```
约束条件为*受扰动的*均值动态
@@ -306,6 +327,12 @@ infl = -intercept * slope / (1 + slope**2)
```
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 主导逃逸路径及其沿路径的通货膨胀率
+ name: fig-esc-dominant-path
+---
fig, axes = plt.subplots(1, 2, figsize=(12, 5))
axes[0].plot(esc.t, intercept, label='intercept $\\gamma_{-1}$')
@@ -353,15 +380,18 @@ R0 = model.M(γ_sce)
R0_inv = np.linalg.inv(R0)
σ1σ2 = model.σ1 * model.σ2
+x_sce = model.x(γ_sce)
+
+# each pair of shock realizations induces its own escape forcing
candidates = {
- "{(1,1),(-1,-1)} → Ramsey": np.array([σ1σ2, 0.0]),
+ "{(1,1),(-1,-1)} → Ramsey": np.array([σ1σ2, 0.0]),
"{(1,-1),(-1,1)} → higher π": np.array([-σ1σ2, 0.0]),
- "{(1,1),(1,-1)}": R0 @ (R0_inv @ np.array([model.x(γ_sce) * model.σ1, model.σ1])),
- "{(-1,1),(-1,-1)}": R0 @ (R0_inv @ np.array([-model.x(γ_sce) * model.σ1, -model.σ1])),
+ "{(1,1),(1,-1)}": np.array([x_sce * model.σ1, model.σ1]),
+ "{(-1,1),(-1,-1)}": np.array([-x_sce * model.σ1, -model.σ1]),
}
for name, force in candidates.items():
- v = R0_inv @ force
+ v = R0_inv @ force # belief velocity along this candidate
print(f" {name:28s} |velocity| = {np.linalg.norm(v):.3f}")
```
@@ -369,7 +399,7 @@ for name, force in candidates.items():
其镜像组合 $\{(1,-1),(-1,1)\}$ 具有相同的速率,但方向错误——指向*更高*的通货膨胀——而在那个方向上均值动态会与之对抗,很快将其拉回。
-因此这场竞速的胜者是朝向拉姆齐方向的路径,它所诱导出的逃逸强制项正是 {eq}`en_force` 中的 $R^{-1}(\sigma_1\sigma_2, 0)'$。
+因此这场竞速的胜者是朝向拉姆齐方向的路径,它所诱导出的逃逸强制项正是 {eq}`en_force` 中的 $R^{-1}(\sigma_1\sigma_2, 0)^\top$。
## 均值动态强化逃逸
@@ -382,6 +412,12 @@ for name, force in candidates.items():
我们可以通过绘制信念空间中的均值动态向量场来观察这一点。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 叠加了逃逸路径的均值动态向量场
+ name: fig-esc-vector-field
+---
gs = np.linspace(-1.2, 0.1, 16) # slope
gi = np.linspace(4.5, 10.5, 16) # intercept
GS, GI = np.meshgrid(gs, gi)
@@ -436,7 +472,13 @@ CWS 的一个引人注目的发现是,当政府的模型更加丰富时,逃
一个更丰富的模型使政府能够检测到自然率假说更为微妙的分布滞后("归纳假说")版本,因此它更容易朝拉姆齐方向逃逸。
```{note}
-逃逸动态继承了使均值动态如此有用的同一种"近似确定性":对于较小的增益,{doc}`phillips_learning` 中的随机模拟会紧贴此处推导出的确定性逃逸路径。下一讲 {doc}`phillips_priors` 表明,政府关于其系数如何漂移的*先验*会重塑这两种动态——甚至可能使逃逸变成一个确定性的*循环*。
+逃逸动态继承了使均值动态如此有用的同一种"近似确定性":
+对于较小的增益,{doc}`phillips_learning` 中的随机模拟会紧贴此处推导出的确定性
+逃逸路径。
+
+下一讲 {doc}`phillips_priors` 表明,政府关于其系数如何漂移的*先验*
+会重塑这两种动态——甚至可能使逃逸变成一个确定性的
+*循环*。
```
## 逃离高波动通货膨胀
@@ -456,7 +498,7 @@ CWS 的一个引人注目的发现是,当政府的模型更加丰富时,逃
```{math}
:label: en_vol
-E(\sigma_\pi \mid \gamma) = \left[ \sigma_2^2 + \left(\frac{\gamma_1}{1 + \gamma_1^2}\right)^2 \sigma_3^2 \right]^{1/2} .
+\mathbb{E}(\sigma_\pi \mid \gamma) = \left[ \sigma_2^2 + \left(\frac{\gamma_1}{1 + \gamma_1^2}\right)^2 \sigma_3^2 \right]^{1/2} .
```
在自我确认均衡 $\gamma_1 = -\theta$ 处,政府相信政策是有效的,并积极对抗 $W_3$,因此通货膨胀是*波动的*。
@@ -466,6 +508,12 @@ E(\sigma_\pi \mid \gamma) = \left[ \sigma_2^2 + \left(\frac{\gamma_1}{1 + \gamma
将 {eq}`en_vol` 应用于我们已经计算出的信念路径,可以看出通货膨胀的水平与波动性*同步*逃逸。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 通货膨胀的水平与波动性同步逃逸
+ name: fig-esc-volatility
+---
σ3 = 0.9 # size of the stabilizable shock
infl_vol = np.sqrt(model.σ2**2 + (slope / (1 + slope**2))**2 * σ3**2)
@@ -495,6 +543,12 @@ plt.show()
如果按字面理解,这意味着一个经济体恰恰在需要稳定化的冲击*较少*时,才更有可能逃向低通胀——这为二十世纪八十年代中期出现的经济平静与伴随而来的反通胀之间提供了一种颇具启发性的联系。
+这同时也是一个关于数据的论断,它指向了本系列讲座压轴的实证内容。
+
+如果经济平静与反通胀是同时到来的,那么一个统计模型必须能够区分*冲击的缩小*与*动态的转变*,才能判断究竟是何者导致了何者。
+
+{doc}`phillips_drifts_volatilities` 构建了一个可以同时容纳这两条渠道的模型,并让数据来判定二者各自的贡献。
+
## 练习
```{exercise-start}
@@ -530,7 +584,7 @@ for σ in [0.2, 0.3, 0.4, 0.5]:
print(f"σ = {σ}: exit time along the escape path = {s.t[-1]:.2f}")
```
-更大的 $\sigma$ 会使逃逸强制项 $R^{-1}(\sigma_1\sigma_2, 0)'$ 变得更强,因此信念沿逃逸路线移动得更快(沿确定性路径的退出*时间*更短)。
+更大的 $\sigma$ 会使逃逸强制项 $R^{-1}(\sigma_1\sigma_2, 0)^\top$ 变得更强,因此信念沿逃逸路线移动得更快(沿确定性路径的退出*时间*更短)。
请注意,这与逃逸的*频率*不同,后者由行动值 $\bar S$ 与增益 $\varepsilon$ 支配;一个噪声更大的经济体,一旦逃逸开始,会更快地走完某条既定的逃逸路线。
@@ -577,4 +631,4 @@ for frac in [0.0, 0.25, 0.5, 0.75, 0.95]:
CWS 所强调的对抗力量,仅局限于均衡附近的一个*极小*邻域——逃逸动态只需要将信念推出这个邻域,之后均值动态便会完成剩余的工作。
```{solution-end}
-```
\ No newline at end of file
+```
diff --git a/lectures/phillips_learning.md b/lectures/phillips_learning.md
index 5ff650d..385f367 100644
--- a/lectures/phillips_learning.md
+++ b/lectures/phillips_learning.md
@@ -46,15 +46,18 @@ translation:
# 适应性学习与逃逸动态
+```{index} single: Phillips Curve; Adaptive Learning and Escape Dynamics
+```
+
```{contents} Contents
:depth: 2
```
## 概述
-本讲座是*菲利普斯曲线权衡*系列讲座的集大成之作。
+本讲座是 {cite}`Sargent1999` 所构建理论的集大成之作。
-它遵循 {cite}`Sargent1999` 的第 8 章,这是该书最具野心的一章。
+它遵循该书第 8 章——这是全书最具野心的一章,而本系列此后的所有内容,要么进一步深化其分析({doc}`phillips_escaping_nash`、{doc}`phillips_priors`),要么将其工具应用于更晚近的一段历史事件({doc}`phillips_lost_conquest`),要么对其核心论断进行实证检验({doc}`phillips_drifts_volatilities`)。
在 {doc}`phillips_self_confirming` 中,政府对菲利普斯曲线持有*固定*的信念——这些信念被由它们本身所生成的数据所证实。
@@ -70,7 +73,7 @@ translation:
答案取决于一个单一的参数——支配旧数据被折扣速度的**增益**(gain):
* 若采用实现最小二乘法的*递减*增益,均值动态会将经济拉向自我确认均衡,我们不会得到什么新结果:系统被困在纳什结果附近。
-* 若采用*常数*增益,主体会对过去数据打折扣,收敛被阻止,**新的结果便会涌现**。系统会反复地从自我确认均衡*逃逸*向拉姆齐(零通胀)结果——这种自发的稳定化现象,颇似沃尔克时代的到来。
+* 若采用*常数*增益,主体会对过去数据打折扣,收敛被阻止,**新的结果便会涌现**:系统会反复地从自我确认均衡*逃逸*向拉姆齐(零通胀)结果,呈现出颇似沃尔克时代到来的自发稳定化现象。
这些逃逸正是 {doc}`phillips_two_stories` 中"econometric 政策评估的证明"故事的核心:一个学习索洛-托宾分布滞后版本自然率假说的适应性政府,会因偶然的观察结果而稳定住通货膨胀。
@@ -98,7 +101,7 @@ from scipy.linalg import solve_discrete_are
在经典识别方案下,自我确认均衡由政府关于某些总体矩及其所隐含的回归系数的信念所确定。
-在经典识别方案下,这些信念由三元组 $(\gamma, \, E X_{C} X_{C}', \, E U X_{C})$ 度量,其中 $\gamma$ 是菲利普斯曲线系数向量。
+在经典识别方案下,这些信念由三元组 $(\gamma, \, E X_{C} X_{C}^\top, \, E U X_{C})$ 度量,其中 $\gamma$ 是菲利普斯曲线系数向量。
在本讲座的适应性模型中,这些对象的*时间 $t$ 的取值*是经济的状态变量之一;它们只有在自我确认均衡中才不再作为状态变量存在,因为在那里它们是常数。
@@ -108,8 +111,8 @@ from scipy.linalg import solve_discrete_are
:label: pl_scemoments
\begin{aligned}
-E\, R_{XC}^{-1}(\gamma)\left[ U_t X_{Ct}' - \left(X_{Ct} X_{Ct}'\right)\gamma \right] &= 0, \\
-E\, X_{Ct} X_{Ct}' - R_{XC}(\gamma) &= 0,
+\mathbb{E}\, R_{XC}^{-1}(\gamma)\left[ U_t X_{Ct}^\top - \left(X_{Ct} X_{Ct}^\top \right)\gamma \right] &= 0, \\
+\mathbb{E}\, X_{Ct} X_{Ct}^\top - R_{XC}(\gamma) &= 0,
\end{aligned}
```
@@ -134,9 +137,9 @@ E\, X_{Ct} X_{Ct}' - R_{XC}(\gamma) &= 0,
```{math}
:label: pl_bdef
-E\left[F(\phi, \zeta)\right] = 0,
+\mathbb{E}\left[F(\phi, \zeta)\right] = 0,
\qquad
-b(\phi) \equiv E\left[F(\phi, \zeta)\right],
+b(\phi) \equiv \mathbb{E}\left[F(\phi, \zeta)\right],
```
其中 $\zeta$ 是一个随机向量,期望是关于其分布计算的(同样,该分布依赖于 $\phi$)。
@@ -165,7 +168,7 @@ b(\phi_f) = 0 .
这正是 {doc}`phillips_self_confirming` 中用来计算自我确认均衡的松弛算法。
-每一步都需要计算数学期望 $b(\phi) = E[F(\phi, \zeta)]$——这正是我们在那里需要矩(李雅普诺夫)公式的原因。
+每一步都需要计算数学期望 $b(\phi) = \mathbb{E}[F(\phi, \zeta)]$——这正是我们在那里需要矩(李雅普诺夫)公式的原因。
### 随机逼近
@@ -194,7 +197,11 @@ t_n = \sum_{k=0}^n a_k ,
增益序列 $\{a_n\}$ 不同的递减速率会产生不同的逼近过程,因为它们改变了从实际时间 $n$ 到人为时间 $t_n$ 的映射 {eq}`pl_artificial`。
```{note}
-递归随机逼近起源于 {cite}`RobbinsMonro1951`,他设计了 {eq}`pl_sa` 用于在噪声观测下寻找回归函数的根;以及 {cite}`KieferWolfowitz1952`,他将其改造用于寻找回归函数的最大值(下文提到的"K-W"算法)。用于分析此类递归的"ODE 方法"——即用微分方程的解来逼近插值过程——归功于 {cite}`Ljung1977`;书籍级的详尽处理见 {cite}`BenvenisteMetivierPriouret1990` 和 {cite}`KushnerYin2003`。将其用于研究自我指涉宏观经济模型中的学习问题,由 {cite}`MarcetSargent1989` 开创,并被 {cite}`EvansHonkapohja2001` 全面发展。
+递归随机逼近起源于 {cite}`RobbinsMonro1951`,他设计了 {eq}`pl_sa` 用于在噪声观测下寻找回归函数的根;以及 {cite}`KieferWolfowitz1952`,他将其改造用于寻找回归函数的最大值(下文提到的“K-W”算法)。
+
+用于分析此类递归的“ODE 方法”——即用微分方程的解来逼近插值过程——归功于 {cite}`Ljung1977`;书籍级的详尽处理见 {cite}`BenvenisteMetivierPriouret1990` 和 {cite}`KushnerYin2003`。
+
+将其用于研究自我指涉宏观经济模型中的学习问题,由 {cite}`MarcetSargent1989` 开创,并被 {cite}`EvansHonkapohja2001` 全面发展。
```
### 均值动态
@@ -263,17 +270,19 @@ t_n = \sum_{k=0}^n a_k ,
```{math}
:label: pl_mgf
-H(\theta, \phi) = \log E \exp\left(\theta' F(\phi, \zeta)\right),
+H(\theta, \phi) = \log \mathbb{E} \exp\left(\theta^\top F(\phi, \zeta)\right),
```
其中期望是关于 $\zeta$ 的分布计算的。
```{note}
-方程 {eq}`pl_mgf` 是一个启发性的简写。实际进入理论的对象是一个*时间平均*极限;{cite}`DupuisKushner1987` 和 {cite}`KushnerYin2003` 假设对每个 $\delta > 0$,以下极限在任意紧集上关于 $\phi_i, \alpha_i$ 一致存在:
+方程 {eq}`pl_mgf` 是一个启发性的简写。
+
+实际进入理论的对象是一个*时间平均*极限;{cite}`DupuisKushner1987` 和 {cite}`KushnerYin2003` 假设对每个 $\delta > 0$,以下极限在任意紧集上关于 $\phi_i, \alpha_i$ 一致存在:
$$
\sum_{i=0}^{T/\delta - 1} \delta\, H(\alpha_i, \phi_i)
= \lim_{N \to \infty} \frac{\delta}{N}
- \log E \exp \sum_{i=0}^{T/\delta - 1} \alpha_i'
+ \log \mathbb{E} \exp \sum_{i=0}^{T/\delta - 1} \alpha_i^\top
\sum_{j=iN}^{iN+N-1} F(\phi_i, \zeta_j) .
$$
内部的求和对长度为 $N$ 的一个区块内的创新项做平均;双重极限使我们能够处理序列相关的创新项。
@@ -284,10 +293,10 @@ $$
```{math}
:label: pl_legendre
-L(\beta, \phi) = \sup_\theta \left[ \theta'\beta - H(\theta, \phi) \right] .
+L(\beta, \phi) = \sup_\theta \left[ \theta^\top \beta - H(\theta, \phi) \right] .
```
-第三是**行动泛函**(action functional),它衡量候选逃逸路径 $\phi(\cdot)$ 的"代价":
+第三是**行动泛函**(action functional),它衡量候选逃逸路径 $\phi(\cdot)$ 的“代价”:
```{math}
:label: pl_action
@@ -326,7 +335,7 @@ A = \left\{ \phi(\cdot) \in C[0,T] : \phi(T) \in \partial D \right\} .
换句话说:*在系统逃离集合 $D$ 的条件下,它离开该集合的位置会接近最小行动路径的终点*——因此逃逸具有确定的方向和形态,尽管它是由偶然性触发的。
-正是在这个意义上,下文中的逃逸虽然不需要任何巨大的冲击触发,却"看起来有目的性":它们遵循 {eq}`pl_escapeproblem` 所规定的最小行动路径。
+正是在这个意义上,下文中的逃逸虽然不需要任何巨大的冲击触发,却“看起来有目的性”:它们遵循 {eq}`pl_escapeproblem` 所规定的最小行动路径。
一个关键的对比:均值动态 {eq}`pl_ode` *不*依赖于其周围的噪声,而逃逸路径*却*依赖于噪声——噪声不仅在 {eq}`pl_ode` 周围添加随机波动,它还开辟出了这第二类路径。
@@ -346,7 +355,7 @@ $$
F(\phi, \zeta) = b(\phi) + \sigma(\phi)\, \zeta,
$$
-其中 $\zeta_n$ 是平稳的且服从高斯分布,但不一定是序列不相关的,并定义 $R = \sum_j E\, \zeta_t \zeta_{t-j}'$。
+其中 $\zeta_n$ 是平稳的且服从高斯分布,但不一定是序列不相关的,并定义 $R = \sum_j \mathbb{E}\, \zeta_t \zeta_{t-j}^\top$。
那么行动泛函取如下二次形式:
@@ -354,22 +363,26 @@ $$
:label: pl_action2
S(T, \phi) = \frac{1}{2} \int_0^T
-\left(\tfrac{d}{ds}\phi - b(\phi)\right)'
-\left[\sigma(\phi)\, R\, \sigma(\phi)'\right]^{+}
+\left(\tfrac{d}{ds}\phi - b(\phi)\right)^\top
+\left[\sigma(\phi)\, R\, \sigma(\phi)^\top \right]^{+}
\left(\tfrac{d}{ds}\phi - b(\phi)\right)
h(s)\, ds ,
```
-其中 $(\cdot)^{+}$ 是摩尔-彭若斯广义逆(用于处理 $\sigma R \sigma'$ 可能出现的随机奇异性)。
+其中 $(\cdot)^{+}$ 是摩尔-彭若斯广义逆(用于处理 $\sigma R \sigma^\top$ 可能出现的随机奇异性)。
权重 $h(s)$ 依赖于增益:当 $a_n = a_0 / n^\gamma$ 中的 $\gamma = 1$ 时,$h(s) = \exp(s)$;当 $\gamma < 1$ 时,$h(s) = 1$。
-将 {eq}`pl_action2` 理解为一种代价,它惩罚*实际漂移* $\tfrac{d}{ds}\phi$ 与*均值漂移* $b(\phi)$ 之间的偏离,并按局部噪声协方差 $\sigma R \sigma'$ 的逆对每个方向加权。
+将 {eq}`pl_action2` 理解为一种代价,它惩罚*实际漂移* $\tfrac{d}{ds}\phi$ 与*均值漂移* $b(\phi)$ 之间的偏离,并按局部噪声协方差 $\sigma R \sigma^\top$ 的逆对每个方向加权。
因此最小行动逃逸路径会穿过均值动态较弱而噪声信息量较大的区域——而在我们的模型中,这恰恰是*归纳假设*所指的方向。
```{note}
-这个二次行动泛函正是 {cite}`ChoWilliamsSargent2002` 中所最小化的对象,该文是本讲座模型的正式发表版本。他们求解了纳什自我确认均衡的控制问题 {eq}`pl_escapeproblem`,并从解析上证明了最小行动逃逸会将通货膨胀权重之和推向激活归纳假设的值——也就是趋向拉姆齐结果。{cite}`SargentWilliams2005` 研究了政府先验(等价地,即增益算法的协方差结构,我们的 $P_0$ 与遗忘因子)如何重塑逃逸,而 {cite}`Kasa2004` 将同样的大偏差机制应用于反复出现的货币危机。
+这个二次行动泛函正是 {cite}`ChoWilliamsSargent2002` 中所最小化的对象,该文是本讲座模型的正式发表版本。
+
+他们求解了纳什自我确认均衡的控制问题 {eq}`pl_escapeproblem`,并从解析上证明了最小行动逃逸会将通货膨胀权重之和推向激活归纳假设的值——也就是趋向拉姆齐结果。
+
+{cite}`SargentWilliams2005` 研究了政府先验(等价地,即增益算法的协方差结构,我们的 $P_0$ 与遗忘因子)如何重塑逃逸,而 {cite}`Kasa2004` 将同样的大偏差机制应用于反复出现的货币危机。
```
### 从计算到适应
@@ -384,7 +397,17 @@ h(s)\, ds ,
2. 递减速度更慢的增益序列——极限情形下即为对过去打折扣的常数增益——则*阻止*了这种拉力,并增加了逃逸动态影响结果的频率。
```{note}
-一段简短的思想史。{cite}`Lucas_Prescott_1971` 曾摒弃对矩条件 {eq}`pl_scezero` 进行迭代这一计算策略,但 {cite}`Townsend1983` 却使用了它。{cite}`Woodford1990` 和 {cite}`MarcetSargent1989` 用均值动态 {eq}`pl_ode` 建立了含有自我指涉的模型中最小二乘学习收敛于理性预期的条件,两者都要求 $b(\phi)$ 的连续性。曹寅坤(In-Koo Cho)研究了带有*不连续* $b(\phi)$ 的问题,这种不连续性源自可信度和搜索问题中不连续的决策规则(触发策略);为了使最小二乘学习逼近理性预期,他使用了满足 $\tfrac{1}{\log n} < a_n < \tfrac{1}{\sqrt n}$ 的增益,这为 {eq}`pl_sa` 产生了一个*扩散*逼近,促进了足够的试探以发现均衡。{cite}`KandoriMailathRob1993` 运用相关数学方法,通过突变在博弈中选择长期均衡,而罗杰·迈尔森(Roger Myerson)将逃逸路径的计算应用于一个投票问题。这些学习方法的现代综合体现在 {cite}`EvansHonkapohja2001` 中。
+一段简短的思想史。
+
+{cite}`Lucas_Prescott_1971` 曾摒弃对矩条件 {eq}`pl_scezero` 进行迭代这一计算策略,但 {cite}`Townsend1983` 却使用了它。
+
+{cite}`Woodford1990` 和 {cite}`MarcetSargent1989` 用均值动态 {eq}`pl_ode` 建立了含有自我指涉的模型中最小二乘学习收敛于理性预期的条件,两者都要求 $b(\phi)$ 的连续性。
+
+曹寅坤(In-Koo Cho)研究了带有*不连续* $b(\phi)$ 的问题,这种不连续性源自可信度和搜索问题中不连续的决策规则(触发策略);为了使最小二乘学习逼近理性预期,他使用了满足 $\tfrac{1}{\log n} < a_n < \tfrac{1}{\sqrt n}$ 的增益,这为 {eq}`pl_sa` 产生了一个*扩散*逼近,促进了足够的试探以发现均衡。
+
+{cite}`KandoriMailathRob1993` 运用相关数学方法,通过突变在博弈中选择长期均衡,而罗杰·迈尔森(Roger Myerson)将逃逸路径的计算应用于一个投票问题。
+
+这些学习方法的现代综合体现在 {cite}`EvansHonkapohja2001` 中。
```
## 适应性模型
@@ -398,9 +421,9 @@ h(s)\, ds ,
```{math}
:label: pl_belief
-U_t = \gamma' X_{C,t} + \varepsilon_{C,t},
+U_t = \gamma^\top X_{C,t} + \varepsilon_{C,t},
\qquad
-X_{C,t} = \begin{bmatrix} y_t & U_{t-1} & U_{t-2} & y_{t-1} & y_{t-2} & 1 \end{bmatrix}' .
+X_{C,t} = \begin{bmatrix} y_t & U_{t-1} & U_{t-2} & y_{t-1} & y_{t-2} & 1 \end{bmatrix}^\top.
```
在时间 $t$ 到达时,凭借估计值 $\gamma_{t-1}$,政府通过求解菲尔普斯问题来设定通货膨胀的系统性部分,*就好像* $\gamma_{t-1}$ 将永远支配菲利普斯曲线一样:
@@ -410,7 +433,7 @@ X_{C,t} = \begin{bmatrix} y_t & U_{t-1} & U_{t-2} & y_{t-1} & y_{t-2} & 1 \end{b
y_t = h(\gamma_{t-1}) X_{t-1} + v_{2t},
\qquad
-X_{t-1} = \begin{bmatrix} U_{t-1} & U_{t-2} & y_{t-1} & y_{t-2} & 1 \end{bmatrix}' .
+X_{t-1} = \begin{bmatrix} U_{t-1} & U_{t-2} & y_{t-1} & y_{t-2} & 1 \end{bmatrix}^\top.
```
随后它通过**递归最小二乘法**(RLS)更新其信念:
@@ -419,8 +442,8 @@ X_{t-1} = \begin{bmatrix} U_{t-1} & U_{t-2} & y_{t-1} & y_{t-2} & 1 \end{bmatrix
:label: pl_rls
\begin{aligned}
-\gamma_t &= \gamma_{t-1} + g_t R_{XC,t}^{-1} X_{C,t}\left(U_t - \gamma_{t-1}' X_{C,t}\right), \\
-R_{XC,t} &= R_{XC,t-1} + g_t\left(X_{C,t} X_{C,t}' - R_{XC,t-1}\right),
+\gamma_t &= \gamma_{t-1} + g_t R_{XC,t}^{-1} X_{C,t}\left(U_t - \gamma_{t-1}^\top X_{C,t}\right), \\
+R_{XC,t} &= R_{XC,t-1} + g_t\left(X_{C,t} X_{C,t}^\top - R_{XC,t-1}\right),
\end{aligned}
```
@@ -440,14 +463,14 @@ $$
给定一个信念 $\gamma$,决策规则 $h(\gamma)$ 求解一个 LQ 控制问题。
-将所信奉的菲利普斯曲线写为 $U_t = \gamma_0 y_t + c' s_t$,其中 $\gamma_0$ 是当期通货膨胀的系数,$c$ 收集了状态 $s_t = X_{t-1}$ 上的系数。
+将所信奉的菲利普斯曲线写为 $U_t = \gamma_0 y_t + c^\top s_t$,其中 $\gamma_0$ 是当期通货膨胀的系数,$c$ 收集了状态 $s_t = X_{t-1}$ 上的系数。
-政府最小化 $E\sum_t \delta^t (U_t^2 + y_t^2)$,因此每期损失为 $s_t' (cc') s_t + (\gamma_0^2 + 1) y_t^2 + 2\gamma_0\, y_t\, c' s_t$,状态按 $s_{t+1} = A s_t + B y_t$ 演化,其中:
+政府最小化 $\mathbb{E}\sum_t \delta^t (U_t^2 + y_t^2)$,因此每期损失为 $s_t^\top (cc^\top) s_t + (\gamma_0^2 + 1) y_t^2 + 2\gamma_0\, y_t\, c^\top s_t$,状态按 $s_{t+1} = A s_t + B y_t$ 演化,其中:
$$
s_{t+1} = \begin{bmatrix} U_t \\ U_{t-1} \\ y_t \\ y_{t-1} \\ 1 \end{bmatrix},
\qquad
-U_t = c' s_t + \gamma_0 y_t .
+U_t = c^\top s_t + \gamma_0 y_t .
$$
我们使用 `scipy` 的离散代数李卡提方程求解器来求解这个折扣 LQ 问题。
@@ -572,6 +595,12 @@ print(f" h(γ_sce) = {np.round(h_sce, 3)} (a constant rule of "
首先,采用最小二乘法(递减增益)。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "最小二乘法下的经典适应性模型:通货膨胀紧贴着自我确认值"
+ name: fig-learn-ls
+---
ls = simulate(model, λ=1.0, T_prior=5000, n=1000, seed=1)
fig, ax = plt.subplots(figsize=(9, 4.5))
@@ -580,7 +609,6 @@ ax.axhline(5, color='k', ls='--', lw=1, label='self-confirming (Nash)')
ax.axhline(0, color='C2', ls=':', lw=1, label='Ramsey')
ax.set_xlabel('$t$')
ax.set_ylabel('inflation $y_t$')
-ax.set_title('Figure 8.1: classical adaptive model, least squares')
ax.legend()
plt.show()
```
@@ -594,6 +622,12 @@ plt.show()
现在赋予政府一个*常数*增益 $\lambda = 0.975$,使其对过去数据打折扣。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "Classical adaptive model under a constant gain: recurrent escapes toward Ramsey"
+ name: fig-learn-cgain
+---
cg = simulate(model, λ=0.975, T_prior=300, n=1000, seed=1)
fig, ax = plt.subplots(figsize=(9, 4.5))
@@ -602,14 +636,10 @@ ax.axhline(5, color='k', ls='--', lw=1, label='self-confirming (Nash)')
ax.axhline(0, color='C2', ls=':', lw=1, label='Ramsey')
ax.set_xlabel('$t$')
ax.set_ylabel('inflation $y_t$')
-ax.set_title('Figure 8.2: classical adaptive model, constant gain '
- r'($\lambda = 0.975$)')
ax.legend()
plt.show()
```
-图景完全不同了。
-
通货膨胀最初接近自我确认值 5,随后几乎跌落至零并停留在那里很长一段时间,然后缓慢地朝 5 迈进,却又再一次被推向零。
将系统拉向自我确认均衡的均值动态,遭到了一种反复出现的力量的对抗,这种力量将通货膨胀推向接近拉姆齐结果的水平。
@@ -630,6 +660,12 @@ print(f"constant-gain inflation: mean {cg['y'].mean():.2f}, "
让我们将通货膨胀与这一权重之和一起绘图。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 稳定化恰好对应着通货膨胀权重之和上升趋近于零
+ name: fig-learn-escape-route
+---
fig, axes = plt.subplots(2, 1, figsize=(9, 7), sharex=True)
axes[0].plot(cg['y'], lw=0.8)
@@ -656,6 +692,12 @@ plt.show()
我们可以通过绘制估计出的菲利普斯曲线中常数项与权重之和的联合路径,直接看出这条逃逸路径。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 信念空间中的逃逸路径,按时间着色
+ name: fig-learn-belief-path
+---
fig, ax = plt.subplots(figsize=(8, 6))
sc = ax.scatter(cg['constant'], cg['sumweights'], c=np.arange(len(cg['y'])),
cmap='viridis', s=6)
@@ -685,6 +727,12 @@ plt.show()
降低 $\delta$ 会提高低通货膨胀插曲期间观测到的通货膨胀率,这与归纳假设下菲尔普斯问题的运作机制是一致的。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 随着政府变得更有耐心,向拉姆齐方向的逃逸变得更加显著
+ name: fig-learn-discount
+---
fig, ax = plt.subplots(figsize=(9, 4.5))
for δ in [0.90, 0.95, 0.98]:
m = AdaptivePhillips(δ=δ)
@@ -694,7 +742,6 @@ ax.axhline(0, color='k', lw=0.5)
ax.set_xlabel('$t$')
ax.set_ylabel('inflation $y_t$')
ax.legend()
-ax.set_title('Escapes toward Ramsey deepen as the government becomes patient')
plt.show()
```
@@ -722,7 +769,13 @@ plt.show()
适应性使得政府的信念成为一种隐藏状态,它给通货膨胀和失业率注入了序列相关性——因此,一个外部预测者若使用随机系数模型,或者进行卢卡斯在其批判 {cite}`lucas1976econometric` 中所指出的那种不断调整,会做得更好。
-在这个意义上,适应性模型包含了证明 econometric 政策评估的基础——这正是 {doc}`phillips_two_stories` 中两个故事的第二个。
+在这个意义上,适应性模型包含了证明经济计量政策评估的基础——这正是 {doc}`phillips_two_stories` 中两个故事的第二个。
+
+最后这一观察是可以检验的,值得在任何人查看数据之前记录下模型的预测。
+
+如果政府的信念确实是一个隐藏的漂移状态,那么对战后通货膨胀与失业率的简化形式描述应当显示出*漂移的系数*,而不仅仅是漂移的冲击方差。
+
+{doc}`phillips_drifts_volatilities` 将恰好这样一个模型拟合到数据上,并给出了结论——其中包括逃逸机制的一个预测,而数据并未能证实这一预测。
## 练习
@@ -730,14 +783,22 @@ plt.show()
:label: pl_ex1
```
-构建**凯恩斯主义**适应性模型,其中政府沿相反方向拟合菲利普斯曲线,将通货膨胀对失业率做回归。
-
-回归量是 $X_{K,t} = \begin{bmatrix} U_t & U_{t-1} & U_{t-2} & y_{t-1} & y_{t-2} & 1 \end{bmatrix}'$,政府在 $y_t = \beta' X_{K,t} + \varepsilon_{K,t}$ 中估计 $\beta$,然后在求解菲尔普斯问题之前将其反转为 $\gamma$。
+以上所有内容都取决于常数增益,因此值得看看其影响究竟有多大。
-不必重新推导所有内容,而是探索*经典*模型对常数增益的敏感性:用 $\lambda \in \{0.99, 0.975, 0.95\}$ 进行模拟,比较通货膨胀朝拉姆齐逃逸的频率。
+用 $\lambda \in \{0.99, 0.975, 0.95\}$ 模拟经典适应性模型,比较通货膨胀朝拉姆齐逃逸的频率。
更大的增益(更小的 $\lambda$,对过去更快地打折扣)如何影响逃逸的频率?
+```{note}
+本练习的一个更具挑战性的版本是构建**凯恩斯主义**适应性模型,其中政府沿相反方向拟合菲利普斯曲线,将通货膨胀对失业率做回归。
+
+回归量是 $X_{K,t} = \begin{bmatrix} U_t & U_{t-1} & U_{t-2} & y_{t-1} & y_{t-2} & 1 \end{bmatrix}^\top$;
+政府在 $y_t = \beta^\top X_{K,t} + \varepsilon_{K,t}$ 中估计 $\beta$,然后使用 {doc}`phillips_self_confirming` 中的反演公式
+$\gamma_1 = \beta_1^{-1}$、$\gamma_{-1} = -\beta_{-1}/\beta_1$ 将其反转为 $\gamma$,之后再求解菲尔普斯问题。
+
+{cite}`Sargent1999` 指出,这一变体不像经典模型那样容易发生逃逸。
+```
+
```{exercise-end}
```
@@ -755,6 +816,7 @@ for λ in [0.99, 0.975, 0.95]:
ax.axhline(0, color='k', lw=0.5)
ax.set_xlabel('$t$')
ax.set_ylabel('inflation $y_t$')
+ax.set_title('Escape frequency by constant gain')
ax.legend()
plt.show()
```
@@ -796,4 +858,4 @@ for T in [500, 2000, 5000]:
更紧的先验则会使增益始终保持较小,因此最小二乘法能可靠地紧贴自我确认均衡。
```{solution-end}
-```
\ No newline at end of file
+```
diff --git a/lectures/phillips_lost_conquest.md b/lectures/phillips_lost_conquest.md
index 274276b..94bc8c3 100644
--- a/lectures/phillips_lost_conquest.md
+++ b/lectures/phillips_lost_conquest.md
@@ -33,6 +33,9 @@ translation:
# 失落的征服:2020 年代的美联储政策
+```{index} single: Phillips Curve; Fed Policy in the 2020s
+```
+
```{contents} Contents
:depth: 2
```
@@ -93,19 +96,22 @@ from scipy.linalg import solve_discrete_are
前两个要素是现代宏观经济学中记录最为详实的事实之一。
-**持续性下降。**
-从 20 世纪 70 年代到 80 年代,通货膨胀具有很强的持续性;一次冲击会使通货膨胀持续数年之久。
+**持续性下降。** 从 20 世纪 70 年代到 80 年代,通货膨胀具有很强的持续性;一次冲击会使通货膨胀持续数年之久。
+
{cite}`CogleySargentConquest2005` 和 {cite}`StockWatson2007` 记录了 20 世纪 80 年代中期之后持续性的显著下降——通货膨胀开始更快地回归目标值。
-**更平坦的菲利普斯曲线。**
-自 20 世纪 90 年代以来,菲利普斯曲线斜率的估计值趋于零——大衰退之后的“消失的反通胀”便是最典型的例子。
+**更平坦的菲利普斯曲线。** 自 20 世纪 90 年代以来,菲利普斯曲线斜率的估计值趋于零——大衰退之后的“消失的反通胀”便是最典型的例子。
+
这两个事实都曾出现在政策制定者的脑海中。
+
前美联储主席珍妮特·耶伦在 2019 年曾指出:“菲利普斯曲线的斜率……自 20 世纪 60 年代以来已显著下降……而且……通货膨胀的持续性已大大降低。”
+
正如 {cite}`Bernanke2022` 所写:“平坦的菲利普斯曲线意味着通货膨胀作为经济过热的指标可靠性降低了,将通货膨胀重新降至目标水平所需付出的失业代价,可能比过去更高。”
-**实时不确定性。**
-{cite}`Orphanides2001` 强调的第三个要素是,产出缺口在*实时*中会被严重误测,尤其是在商业周期的转折点。
+**实时不确定性。** {cite}`Orphanides2001` 强调的第三个要素是,产出缺口在*实时*中会被严重误测,尤其是在商业周期的转折点。
+
在 2020 年至 2023 年间,实时缺口持续*低于*后来修订后的度量值,因此美联储感知到了更多的经济松弛——这强化了通货膨胀会自行消退的信念。
+
我们在下文使用当前版本的数据,并在结论部分回到实时数据的区别问题上。
## 美联储的漂移系数信念
@@ -123,12 +129,12 @@ from scipy.linalg import solve_discrete_are
美联储通过常增益递归最小二乘法更新 $\theta_t = (\alpha_{0,t}, \rho_t, \kappa_t)$——这正是 {doc}`phillips_learning` 和 {doc}`phillips_priors` 中的算法,增益 $\gamma$ 对过去的数据进行折扣,从而使估计值能够*追踪*漂移:
$$
-\theta_{t+1} = \theta_t + \gamma R_t^{-1} X_t\left(\pi_t - X_t'\theta_t\right),
+\theta_{t+1} = \theta_t + \gamma R_t^{-1} X_t\left(\pi_t - X_t^\top \theta_t\right),
\qquad
-R_{t+1} = R_t + \gamma\left(X_t X_t' - R_t\right),
+R_{t+1} = R_t + \gamma\left(X_t X_t^\top - R_t\right),
$$
-其中 $X_t = (1, \pi_{t-1}, x_t)'$。
+其中 $X_t = (1, \pi_{t-1}, x_t)^\top$。
我们从 FRED 下载季度个人消费支出(PCE)通货膨胀率、国会预算办公室(CBO)产出缺口和联邦基金利率。
@@ -172,6 +178,12 @@ beliefs = estimate_beliefs(data)
```
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 美联储感知的通货膨胀持续性和菲利普斯曲线斜率
+ name: fig-lc-beliefs
+---
fig, axes = plt.subplots(2, 1, figsize=(10, 6), sharex=True)
axes[0].plot(beliefs['rho'])
axes[0].axhline(1, color='k', lw=0.5, ls=':')
@@ -188,8 +200,6 @@ plt.tight_layout()
plt.show()
```
-这两幅图讲述了这个故事。
-
在通货膨胀高企的 20 世纪 70 年代和 80 年代,感知的**持续性** $\rho_t$ 接近 1,随后在 80 年代中期之后逐渐下降,在 2008 年之后触底——然后在 2022 年信念更新恢复时*又跳回*接近 1 的水平,恰好也是美联储放弃“暂时性”表述并开始收紧政策的时候。
感知的**斜率** $\kappa_t$ 在整个 2010 年代趋向于零:菲利普斯曲线趋于平坦。
@@ -200,7 +210,7 @@ plt.show()
美联储在每一期通过求解一个线性二次型菲尔普斯问题来设定其政策利率,将其*当前*的估计值视为将永远成立——这正是我们在 {doc}`phillips_learning` 中遇到的 {cite}`Kreps1998` 的预期效用假设。
-将信念菲利普斯曲线 {eq}`lc_pc` 与一条固定的“IS 曲线” $x_t = b_0 + b_1 x_{t-1} + g(i_{t-1} - \pi_{t-1}) + \varepsilon^x_t$ 相配对,可以得到状态 $X_t = (1, \pi_t, x_t, i_{t-1})'$ 的线性动态:
+将信念菲利普斯曲线 {eq}`lc_pc` 与一条固定的“IS 曲线” $x_t = b_0 + b_1 x_{t-1} + g(i_{t-1} - \pi_{t-1}) + \varepsilon^x_t$ 相配对,可以得到状态 $X_t = (1, \pi_t, x_t, i_{t-1})^\top$ 的线性动态:
$$
X_{t+1} = A_t X_t + B_t\, i_t + C \varepsilon_{t+1},
@@ -230,7 +240,7 @@ b0, b1, g = np.linalg.lstsq(X_is, x[1:], rcond=None)[0]
β, π_star, λ_x, η = 0.95, 2.0, 0.2, 0.5
-def phelps_rate(θ, state):
+def phelps_rate(θ, state, η=η):
"在给定信念 θ=(α₀,ρ,κ) 和状态的情况下计算(主观上)最优的联邦基金利率。"
α0, ρ, κ = θ
A = np.array([[1, 0, 0, 0],
@@ -270,13 +280,18 @@ optimal = pd.Series(opt, index=data.index[1:])
```
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 信念驱动的菲尔普斯规则与实际联邦基金利率对比
+ name: fig-lc-phelps-rate
+---
fig, ax = plt.subplots(figsize=(10, 4.5))
window = slice('1991', None)
ax.plot(optimal[window], 'C0', label="菲尔普斯问题的建议利率")
ax.plot(data['i'][window], 'C3', lw=1, label='实际联邦基金利率')
ax.set_xlabel('年份')
ax.set_ylabel('百分比')
-ax.set_title("信念驱动的菲尔普斯规则与实际政策对比")
ax.legend()
plt.show()
@@ -293,6 +308,12 @@ print(f"1991-2025 年建议利率与实际利率的相关性:{corr:.2f}")
为了分离出漂移信念所起的作用,我们重新计算菲尔普斯建议利率,这次将信念*固定*在 2000 年 1 月的取值上——当时人们仍然认为通货膨胀具有持续性,且菲利普斯曲线更陡峭。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "反事实分析:一个自 2000 年以来未曾更新信念的美联储"
+ name: fig-lc-counterfactual
+---
counterfactual = pd.Series(
[phelps_rate(θ_2000, np.array([1.0, pi[t], x[t], i_[t - 1]]))
for t in range(1, n)],
@@ -305,13 +326,10 @@ ax.plot(counterfactual[w], 'C1--', label='反事实情形(信念冻结于 2000
ax.plot(data['i'][w], 'C3', lw=1, alpha=0.7, label='实际联邦基金利率')
ax.set_xlabel('年份')
ax.set_ylabel('百分比')
-ax.set_title('反事实分析:一个自 2000 年以来未曾更新信念的美联储')
ax.legend()
plt.show()
```
-对比结果十分鲜明。
-
一个持有 2000 年信念的美联储——认为通货膨胀具有持续性且菲利普斯曲线更陡峭——本应在 2021 年*立即而剧烈地*收紧政策,在实际美联储尚未采取任何行动之前,就将联邦基金利率推高至远超 4% 的水平。
这种温和而滞后的反应,并非目标的改变,而是*信念*的改变。
@@ -357,6 +375,12 @@ i_t &= \phi_\pi\, \pi_t ,
让我们通过求解三次方程的稳定根来重现 {prf:ref}`lc_prop1`。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 激进的政策使通货膨胀看起来持续性更低
+ name: fig-lc-persistence
+---
def measured_persistence(φ_π, β=0.99, γ_b=0.5, κ=0.1, σ=1.0):
"稳定的最小状态变量根 λ(φ_π):计量经济学家将测得的持续性。"
coeffs = [β, -(1 + β + κ * σ), 1 + γ_b + κ * σ * φ_π, -γ_b]
@@ -369,10 +393,9 @@ def measured_persistence(φ_π, β=0.99, γ_b=0.5, κ=0.1, σ=1.0):
λ_path = [measured_persistence(φ) for φ in φ_grid]
fig, ax = plt.subplots(figsize=(8, 4.5))
-ax.plot(φ_grid, λ_path)
+ax.plot(φ_grid, λ_path, lw=2)
ax.set_xlabel(r'泰勒规则激进程度 $\phi_\pi$')
ax.set_ylabel(r'测量的持续性 $\lambda$')
-ax.set_title('激进的政策使通货膨胀看起来持续性更低')
plt.show()
```
@@ -399,9 +422,17 @@ plt.show()
这是《征服》一书中反复出现的动态在更高层次上的一次现代重演:这一机制现在通过菲利普斯曲线感知的斜率和持续性,以及政策利率工具发挥作用,而误设并非关于预期,而是关于美联储将简化形式菲利普斯曲线视为结构性的这一*政策内生性*问题——正如 {doc}`phillips_two_stories` 中所描述的“平反”故事那样,忽视了卢卡斯批判。
```{note}
-正如 {cite}`SargentWilliams2025` 所指出的,漂移系数模型纯粹是一个描述性的“开普勒阶段”模型,而非结构性的“牛顿阶段”模型。原论文也承认了另一种解读方式,即 2020 年代的宽松政策起源于财政因素——参见其所引用的财政理论论述——这是对同一政策路径的一种截然不同的解释。
+正如 {cite}`SargentWilliams2025` 所指出的,漂移系数模型纯粹是一个描述性的“开普勒阶段”模型,
+而非结构性的“牛顿阶段”模型。
+
+原论文也承认了另一种解读方式,即 2020 年代的宽松政策起源于财政因素——参见其所引用的财政理论
+论述——这是对同一政策路径的一种截然不同的解释。
```
+本讲将漂移的信念归结于美联储,然后追问这些信念会推荐怎样的政策。
+
+本系列的收尾讲座 {doc}`phillips_drifts_volatilities` 则从相反的方向审视同一时期:它追问数据,经济的*简化形式*究竟是否真的发生了漂移,以及那些看似信念漂移的现象,究竟有多少实际上是冲击方差在漂移。
+
## 练习
```{exercise-start}
@@ -422,10 +453,8 @@ plt.show()
```
```{code-cell} ipython3
-def recommend(θ_path_fn, η_val):
- global η
- η_save = η
- η = η_val
+def recommend(η_val):
+ "在给定平滑权重下,沿整个样本计算菲尔普斯建议利率。"
θ, R = np.array([0.5, 0.9, 0.05]), np.diag([1.0, 10.0, 5.0])
out = []
for t in range(1, n):
@@ -433,17 +462,18 @@ def recommend(θ_path_fn, η_val):
X = np.array([1.0, pi[t - 1], x[t]])
R = R + g_t * (np.outer(X, X) - R)
θ = θ + g_t * np.linalg.solve(R, X * (pi[t] - X @ θ))
- out.append(phelps_rate(θ, np.array([1.0, pi[t], x[t], i_[t - 1]])))
- η = η_save
+ out.append(phelps_rate(θ, np.array([1.0, pi[t], x[t], i_[t - 1]]),
+ η=η_val))
return pd.Series(out, index=data.index[1:])
fig, ax = plt.subplots(figsize=(10, 4.5))
w = slice('2015', None)
for η_val in [0.1, 0.5, 2.0]:
- ax.plot(recommend(None, η_val)[w], lw=1, label=rf'$\eta = {η_val}$')
+ ax.plot(recommend(η_val)[w], lw=1, label=rf'$\eta = {η_val}$')
ax.plot(data['i'][w], 'k:', lw=1.5, label='实际利率')
ax.set_xlabel('年份')
ax.set_ylabel('百分比')
+ax.set_title('不同平滑权重下的菲尔普斯建议利率')
ax.legend()
plt.show()
```
diff --git a/lectures/phillips_misspecified.md b/lectures/phillips_misspecified.md
index cafacd8..f9db70a 100644
--- a/lectures/phillips_misspecified.md
+++ b/lectures/phillips_misspecified.md
@@ -34,6 +34,9 @@ translation:
# 最优错误设定信念
+```{index} single: Phillips Curve; Optimal Misspecified Beliefs
+```
+
```{contents} Contents
:depth: 2
```
@@ -44,6 +47,12 @@ translation:
内容遵循 {cite}`Sargent1999` 的第 6 章。
+在 {doc}`phillips_adaptive` 中,公众使用一个固定的适应性规则来预测通货膨胀,其中的参数 $\lambda$ 是我们直接选定的。
+
+从某种意义上说,这种做法并不令人满意,而 {doc}`phillips_two_stories` 中的*印证*故事是无法容忍这一点的:描述预期的自由参数正是理性预期本应消除的东西。
+
+在这里,我们朝着"赢得"这个参数迈出了第一步——让代理人自己选择该参数,以拟合由他们自身信念所产生的数据。
+
我们描述了贯穿本系列讲座的三个概念性问题:
1. 如何构建这样一种均衡——其中代理人共享一个共同的*错误设定的*最小二乘预测模型,
@@ -283,14 +292,20 @@ print(f"actual one-step forecast error std σ̄_ε = {σ_bar:.4f}")
对于均衡 $C$,我们绘制真实模型与近似模型的均衡谱密度。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 真实价格过程与代理人近似模型的谱密度
+ name: fig-mis-spectra
+---
F = bray.true_spectrum(C_star)
σ_ε2 = fitted_sigma2(bray, C_star, c_star)
G = bray.approx_spectrum(c_star, σ_ε2)
half = bray.N // 2
fig, ax = plt.subplots(figsize=(8, 5))
-ax.plot(bray.ω[:half], np.log(F[:half]), 'C0', label='true model')
-ax.plot(bray.ω[:half], np.log(G[:half]), 'C1--', label='forecasting model')
+ax.plot(bray.ω[:half], np.log(F[:half]), 'C0', label='true model', lw=2)
+ax.plot(bray.ω[:half], np.log(G[:half]), 'C1--', label='forecasting model', lw=2)
ax.set_xlabel(r'angular frequency $\omega$')
ax.set_ylabel('log spectral density')
ax.legend()
@@ -308,6 +323,12 @@ plt.show()
我们通过向每个移动平均表示输入一个单位冲击,来比较两个模型的脉冲响应函数。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 真实模型与近似模型的脉冲响应
+ name: fig-mis-irf
+---
def impulse_response(num_roots, den_roots, T=25):
"IRF of (1 - num L)/(1 - den L): coefficients of the ratio of lag polys."
h = np.empty(T)
@@ -323,8 +344,8 @@ irf_true = scale * impulse_response(1 - C_star, φ) # f(L)
irf_approx = impulse_response(1 - c_star, bray.ρ) # g(L)
fig, ax = plt.subplots(figsize=(8, 5))
-ax.plot(irf_true, 'C0o-', ms=4, label='true model')
-ax.plot(irf_approx, 'C1s--', ms=4, label='approximating model')
+ax.plot(irf_true, 'C0o-', ms=4, label='true model', lw=2)
+ax.plot(irf_approx, 'C1s--', ms=4, label='approximating model', lw=2)
ax.set_xlabel('lag')
ax.set_ylabel('response')
ax.legend()
@@ -373,9 +394,10 @@ b_grid = np.arange(0.1, 0.85, 0.1)
C_of_b = [solve_equilibrium(BrayModel(a=1.0, b=b, σ_u=1.0)) for b in b_grid]
fig, ax = plt.subplots(figsize=(8, 4.5))
-ax.plot(b_grid, C_of_b, 'o-')
+ax.plot(b_grid, C_of_b, 'o-', lw=2)
ax.set_xlabel('feedback parameter $b$')
ax.set_ylabel('equilibrium belief $C$')
+ax.set_title('Equilibrium belief by expectational feedback')
plt.show()
```
@@ -409,6 +431,7 @@ ax.annotate('equilibrium', (C_star, C_star),
(C_star + 0.05, C_star - 0.03))
ax.set_xlabel('$C$')
ax.set_ylabel('$B(C)$')
+ax.set_title('The best-estimate map and its fixed point')
ax.legend()
plt.show()
```
@@ -416,4 +439,4 @@ plt.show()
最优估计映射在均衡信念处与 45 度线相交,证实了 $C = B(C)$。
```{solution-end}
-```
\ No newline at end of file
+```
diff --git a/lectures/phillips_priors.md b/lectures/phillips_priors.md
index 109d460..f73f5e3 100644
--- a/lectures/phillips_priors.md
+++ b/lectures/phillips_priors.md
@@ -36,6 +36,9 @@ translation:
# 先验、逃逸与学习循环
+```{index} single: Phillips Curve; Priors and Learning Cycles
+```
+
```{contents} Contents
:depth: 2
```
@@ -92,7 +95,7 @@ U_n &= u - (\pi_n - \hat x_n) + \sigma_1 W_{1n}, \qquad u > 0, \\
\end{aligned}
```
-其中 $U_n$ 是失业率,$\pi_n$ 是通货膨胀率,$x_n$ 是政府设定的通胀的系统性部分,$\hat x_n$ 是公众的(理性)预测,$W_n = (W_{1n}, W_{2n})'$ 是独立同分布的标准高斯噪声。
+其中 $U_n$ 是失业率,$\pi_n$ 是通货膨胀率,$x_n$ 是政府设定的通胀的系统性部分,$\hat x_n$ 是公众的(理性)预测,$W_n = (W_{1n}, W_{2n})^\top$ 是独立同分布的标准高斯噪声。
由于 $\pi_n - \hat x_n = \sigma_2 W_{2n}$,真实的失业率是 $U_n = u - \sigma_2 W_{2n} + \sigma_1 W_{1n}$——无论系统性政策如何,它都围绕自然率 $u$ 波动。
@@ -108,7 +111,20 @@ U_n = a + b\, \pi_n + \eta_n ,
信念向量为 $\gamma = (a, b)$(截距和斜率),并将 $\eta_n$ 视为外生冲击。
-在相信 {eq}`pp_belief` 的前提下,政府求解菲尔普斯问题——最小化 $\hat E \sum_n \delta^n (U_n^2 + \pi_n^2)$——其静态最优反应将通胀设定为常数
+```{warning}
+{doc}`phillips_escaping_nash` 研究的是同一个静态模型,但信念向量的排列顺序相反,写作
+$\gamma = (\gamma_1, \gamma_{-1})$,*斜率*在前,回归量为 $\Phi = (\pi, 1)$,遵循
+{cite}`ChoWilliamsSargent2002` 的做法。
+
+在这里,我们把*截距*放在前面,$\Phi = (1, \pi)$,遵循 {cite}`SargentWilliams2005` 的做法。
+
+这两者其实是同一个模型在不同坐标下的表示:$a = \gamma_{-1}$ 且 $b = \gamma_1$,因此本讲中的自我确认信念
+$(2u, -1)$ 对应于上一讲中的 $(-1, u(1+\theta^2))$,其中 $\theta = 1$。
+
+诸如 $M$、$V$ 和 $P$ 这样的矩阵,其行和列也相应地进行了转置。
+```
+
+在相信 {eq}`pp_belief` 的前提下,政府求解菲尔普斯问题——最小化 $\hat{\mathbb{E}} \sum_n \delta^n (U_n^2 + \pi_n^2)$——其静态最优反应将通胀设定为常数
```{math}
:label: pp_bestresp
@@ -144,7 +160,7 @@ class StaticPhillips:
有三个信念向量值得命名。
-* **信念 1(纳什):** $b = -1$,截距使政府设定 $x = u$。这是 {cite}`KydlandPrescott1977` 的时间一致性结果。
+* **信念 1(纳什):** $b = -1$,截距使政府设定 $x = u$,这是 {cite}`KydlandPrescott1977` 的时间一致性结果。
* **信念 2(拉姆齐):** $b = 0$,因此政府认为*不存在*权衡,设定 $x = 0$。
* **信念 3(归纳):** 在动态版本中,当前和滞后通胀的系数之和为零,这对于一个有耐心的政府来说也会使通胀趋向 $0$。
@@ -154,7 +170,7 @@ class StaticPhillips:
对于静态模型,这可以手工轻松求解。
-斜率为 $b = \operatorname{cov}(U, \pi)/\operatorname{var}(\pi) = -\sigma_2^2/\sigma_2^2 = -1$,均值匹配给出截距 $a = u + x(\bar\gamma)$。
+斜率为 $b = \operatorname{cov}(U, \pi)/\mathbb{V}[\pi] = -\sigma_2^2/\sigma_2^2 = -1$,均值匹配给出截距 $a = u + x(\bar\gamma)$。
将最优反应 {eq}`pp_bestresp`($b = -1$)代入,得到 $x = a/2$,因此 $a = u + a/2$,即 $a = 2u$。
@@ -189,13 +205,13 @@ print(f"check g_bar(γ_sce) = {model.g_bar(γ_sce)}")
协方差矩阵 $V$ 是政府对**参数漂移的先验信念**——这是我们要放开的对象。
-对于回归量 $\Phi_n = (1, \pi_n)'$,卡尔曼滤波的大样本近似(见 {cite}`BenvenisteMetivierPriouret1990`)为
+对于回归量 $\Phi_n = (1, \pi_n)^\top$,卡尔曼滤波的大样本近似(见 {cite}`BenvenisteMetivierPriouret1990`)为
```{math}
:label: pp_kalman
\begin{aligned}
-\gamma_{n+1} &= \gamma_n + P_n \Phi_n\left(U_n - \Phi_n' \gamma_n\right), \\
+\gamma_{n+1} &= \gamma_n + P_n \Phi_n\left(U_n - \Phi_n^\top \gamma_n\right), \\
P_{n+1} &= P_n - P_n M(\gamma_n) P_n + \sigma^{-2} V ,
\end{aligned}
```
@@ -301,6 +317,12 @@ V(\lambda) = \begin{bmatrix} V^*_{11} & \sqrt\lambda\, V^*_{12} \\ \sqrt\lambda\
对每个 $\lambda$,我们求解里卡蒂方程并观察 $\bar P(\lambda)\, \partial\bar g/\partial\gamma$ 特征值中最大实部:当该值为正时,自证均衡是不稳定的。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 随着斜率先验收紧,自证均衡的稳定性变化
+ name: fig-pri-stability
+---
def V_tighten_slope(λ, V_star):
V = V_star.copy()
V[0, 1] = V[1, 0] = np.sqrt(λ) * V_star[0, 1]
@@ -320,7 +342,6 @@ ax.plot(λ_grid, max_re)
ax.axhline(0, color='k', lw=0.8)
ax.set_xlabel(r'prior-tightening parameter $\lambda$')
ax.set_ylabel('max real part of eigenvalue')
-ax.set_title(r'Figure 4: stability of the SCE as the slope prior tightens')
plt.show()
```
@@ -348,16 +369,22 @@ mask = sol.t > sol.t[-1] - 800
```
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 系数循环,以及信念空间中的极限循环
+ name: fig-pri-cycle
+---
fig, axes = plt.subplots(1, 2, figsize=(12, 5))
-axes[0].plot(sol.t[mask], a_path[mask], label='intercept')
-axes[0].plot(sol.t[mask], b_path[mask], label='slope')
+axes[0].plot(sol.t[mask], a_path[mask], label='intercept', lw=2)
+axes[0].plot(sol.t[mask], b_path[mask], label='slope', lw=2)
axes[0].set_xlabel('time')
axes[0].set_ylabel('coefficient')
axes[0].set_title('Figure 5a: coefficients cycle')
axes[0].legend()
-axes[1].plot(a_path[mask], b_path[mask])
+axes[1].plot(a_path[mask], b_path[mask], lw=2)
axes[1].plot(*γ_sce, 'kx', ms=10, label='SCE')
axes[1].set_xlabel('intercept')
axes[1].set_ylabel('slope')
@@ -373,13 +400,18 @@ plt.show()
由于通胀通过最优反应 {eq}`pp_bestresp` 是系数的函数,信念中的循环表现为通胀在纳什和拉姆齐结果之间振荡的循环。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 沿着学习循环,通胀在纳什和拉姆齐结果之间振荡
+ name: fig-pri-inflation-cycle
+---
fig, ax = plt.subplots(figsize=(9, 4.5))
ax.plot(sol.t[mask], x_path[mask])
ax.axhline(model.u, color='k', ls='--', lw=1, label='Nash')
ax.axhline(0, color='C2', ls=':', lw=1, label='Ramsey')
ax.set_xlabel('time')
ax.set_ylabel('inflation $x$')
-ax.set_title('Figure 6: inflation oscillates between Nash and Ramsey along the cycle')
ax.legend()
plt.show()
```
@@ -431,7 +463,30 @@ print(f"terminal beliefs ≈ {terminal.round(2)} (Belief 2 = [u, 0] = Ramsey)
西姆斯则使用了 $\sigma \neq \sigma_1$(且没有缩小增益),这*错误地分配*了观测到的变异,从而产生了长期、也许是永久性的偏离自证均衡的情形。
-让我们在两种设定下模拟静态模型。
+在我们模拟任何东西之前,这一机制在里卡蒂方程 {eq}`pp_riccati` 中就已经显现出来。
+
+一个对其回归误差赋予*过少*方差的政府,必须将它所观察到的变异归因于其他因素,而唯一可用的其他因素就是其自身系数的漂移。
+
+因此 $\sigma < \sigma_1$ 会使稳态的 $P$ 膨胀,而 $P$ 正是在 {eq}`pp_kalman` 中乘以每个预测误差的量。
+
+因此,低估误差方差会使政府*学习得更快*——而这恰恰削弱了均值动态的拉力。
+
+```{code-cell} ipython3
+def effective_gain(model, σ_govt, ε, λ=1.0):
+ "Steady-state weight the filter puts on a single forecast error at the SCE."
+ V = ε**2 * V_tighten_slope(λ, V_star)
+ P = solve_riccati(V, M_sce, σ_govt)
+ Φ = np.array([1.0, model.x(γ_sce)])
+ return (P @ Φ)[0] / (σ_govt**2 + Φ @ P @ Φ)
+
+ε_common = 0.0002
+for σ_govt, tag in ((model.σ1, 'σ = σ1 (correct) '), (0.1, 'σ ≠ σ1 (Sims) ')):
+ print(f"{tag}: effective gain = {effective_gain(model, σ_govt, ε_common):.5f}")
+```
+
+即使先验 $V$ 和尺度 $\varepsilon$ 在两次运行中都相同,低估误差方差也会使有效增益放大二十多倍。
+
+现在让我们在两种设定下模拟静态模型,保持 $\varepsilon$ 固定不变,这样两者*唯一*的差异就在于政府对其自身回归误差所赋予的方差。
```{code-cell} ipython3
def simulate(model, σ_govt, ε, λ=1.0, T=3000, seed=0):
@@ -456,35 +511,66 @@ def simulate(model, σ_govt, ε, λ=1.0, T=3000, seed=0):
infl[n] = π
return infl
-x_base = simulate(model, σ_govt=model.σ1, ε=0.05, seed=1) # σ = σ1
-x_sims = simulate(model, σ_govt=0.1, ε=0.20, seed=1) # σ ≠ σ1 (Sims-like)
+T_sim, seeds = 6000, range(10)
-print(f"σ = σ1 : mean inflation {x_base.mean():.2f}, "
- f"fraction near Ramsey {(x_base < 2).mean():.0%}")
-print(f"σ ≠ σ1 : mean inflation {x_sims.mean():.2f}, "
- f"fraction near Ramsey {(x_sims < 2).mean():.0%}")
+def summarize(σ_govt):
+ "Mean inflation and time near Ramsey, across seeds."
+ paths = [simulate(model, σ_govt=σ_govt, ε=ε_common, T=T_sim, seed=s)
+ for s in seeds]
+ return (np.array([p.mean() for p in paths]),
+ np.array([(p < 2).mean() for p in paths]))
+
+mean_base, ramsey_base = summarize(model.σ1)
+mean_sims, ramsey_sims = summarize(0.1)
+
+print(f"{'seed':>4} | {'σ = σ1: mean π':>15} {'near Ramsey':>12}"
+ f" | {'σ ≠ σ1: mean π':>15} {'near Ramsey':>12}")
+for s in seeds:
+ print(f"{s:>4} | {mean_base[s]:>15.2f} {ramsey_base[s]:>11.0%}"
+ f" | {mean_sims[s]:>15.2f} {ramsey_sims[s]:>11.0%}")
+print(f"\nnever escaped in {T_sim} periods: "
+ f"σ = σ1: {np.sum(ramsey_base < 0.01)}/{len(seeds)} paths, "
+ f"σ ≠ σ1: {np.sum(ramsey_sims < 0.01)}/{len(seeds)} paths")
```
+这两列的表现相当不同,而这种差异关乎经济*多频繁地*离开纳什水平,而不在于它逃逸后去往何处。
+
+在误差方差设定正确的情况下,逃逸是一个真正罕见的事件:若干条样本路径在六千个周期内始终停留在纳什利率附近,从未逃逸;其余路径则在某个时刻逃逸,此后又被拉回到拉姆齐水平附近。
+
+而在西姆斯的错误分配下,*每条*路径都会逃逸,并且每一条都将大部分时间停留在拉姆齐水平附近。
+
+三个种子无需任何挑选就能说明问题:种子 0 从未逃逸,种子 4 反复逃逸并每次都被拉回,种子 6 逃逸后则长时间远离纳什水平。
+
```{code-cell} ipython3
-fig, axes = plt.subplots(2, 1, figsize=(9, 7), sharex=True)
-axes[0].plot(x_base, lw=0.6)
-axes[0].axhline(model.u, color='k', ls='--', lw=1)
-axes[0].set_ylabel('inflation')
-axes[0].set_title(r'$\sigma = \sigma_1$: recurrent escapes, pulled back to Nash')
-
-axes[1].plot(x_sims, lw=0.6, color='C1')
-axes[1].axhline(model.u, color='k', ls='--', lw=1)
+---
+mystnb:
+ figure:
+ caption: 在误差方差设定正确的情况下逃逸是罕见的,而在西姆斯的错误分配下逃逸则普遍存在
+ name: fig-pri-sims
+---
+show = (0, 4, 6)
+
+fig, axes = plt.subplots(2, 1, figsize=(9.5, 7), sharex=True, sharey=True)
+for s in show:
+ axes[0].plot(simulate(model, σ_govt=model.σ1, ε=ε_common, T=T_sim, seed=s),
+ lw=0.5, label=f'seed {s}')
+ axes[1].plot(simulate(model, σ_govt=0.1, ε=ε_common, T=T_sim, seed=s),
+ lw=0.5, label=f'seed {s}')
+for ax, title in zip(axes, [r'$\sigma = \sigma_1$ (correct): escapes are rare events',
+ r'$\sigma \neq \sigma_1$ (Sims): every path escapes and lingers']):
+ ax.axhline(model.u, color='k', ls='--', lw=1)
+ ax.axhline(0, color='C2', ls=':', lw=1)
+ ax.set_ylabel('inflation')
+ ax.set_title(title)
+ ax.legend(frameon=False, fontsize=8, ncol=3, loc='lower right')
axes[1].set_xlabel('$n$')
-axes[1].set_ylabel('inflation')
-axes[1].set_title(r'$\sigma \neq \sigma_1$ (Sims): prolonged spells near Ramsey')
-
plt.tight_layout()
plt.show()
```
-在误差方差设定正确的情况下,均值动态会重新发挥作用,通胀被反复拉回纳什水平附近。
+在误差方差设定正确的情况下,均值动态占主导地位:经济停留在纳什水平,只有异常连续的一连串冲击才能将其撼动——这正是 {doc}`phillips_learning` 和 {doc}`phillips_escaping_nash` 中呈现的罕见逃逸图景。
-而在西姆斯的错误分配下,这种拉力被削弱,经济则长期停留在拉姆齐结果附近——政府的行为就好像它已经*永久*学会了一个足够好的自然率假说版本。
+而在西姆斯的错误分配下,政府学习得太快,以至于均值动态无法将其维持住,于是它几乎立即逃逸,其行为就好像已经*永久*学会了一个足够好的自然率假说版本。
正如 {cite}`SargentWilliams2005` 所言,这种差异可以用两种等价的方式来理解:要么是西姆斯允许了过多的参数漂移以至于无法收敛,要么是他没有让政府对其回归误差归因足够多的变异。
@@ -520,7 +606,9 @@ plt.show()
它决定了经济是收敛到纳什状态、在纳什和拉姆齐之间循环,还是逃逸到拉姆齐状态并停留在那里。
-最后一讲,{doc}`phillips_lost_conquest`,将同样的工具——固定增益学习、预期效用菲尔普斯问题以及自证均衡——带入当下,用以解释美联储对 2020 年代通胀的应对。
+{doc}`phillips_lost_conquest` 将同样的工具——固定增益学习、预期效用菲尔普斯问题以及自证均衡——带入当下,用以解释美联储对 2020 年代通胀的应对。
+
+{doc}`phillips_drifts_volatilities` 随后为本系列画上句点,对这整套论述进行了一次实证检验:将一个带漂移系数、随机波动率的向量自回归模型拟合到战后数据,探究大通胀在多大程度上源于信念漂移,又在多大程度上仅仅是运气不佳。
## 练习
@@ -566,6 +654,7 @@ ax.plot(λ_grid, max_re, ls='--', label='tighten slope (for comparison)')
ax.axhline(0, color='k', lw=0.8)
ax.set_xlabel(r'$\lambda$')
ax.set_ylabel('max real part of eigenvalue')
+ax.set_title('Stability under a tighter intercept prior')
ax.legend()
plt.show()
```
@@ -609,9 +698,11 @@ for name, V in [("baseline V*", V_star),
print(f"{name:24s}: direction {v.round(3)}, terminal {term.round(2)}")
```
-两种先验都将信念导向斜率为 $0$ 的终端点——即拉姆齐信念——但沿着不同的方向,并到达略有不同的截距。
+两种先验都将信念导向斜率为 $0$ 的终端点——即拉姆齐信念——但沿着不同的方向,且二者到达的截距差异相当大:斜率收紧先验落在的截距略高于基准情形截距的一半。
-目的地(零通胀)是一个稳健的特征;而*路径*则取决于先验的形状。
+在对政策有意义的层面上,目的地是稳健的:斜率为 $0$ 意味着政府认为不存在可利用的权衡关系,此时 {eq}`pp_bestresp` 无论截距如何,都会将通胀设为零。
+
+而*路径*,以及逃逸在拉姆齐线上具体终止于何处,则取决于先验的形状。
```{solution-end}
-```
\ No newline at end of file
+```
diff --git a/lectures/phillips_self_confirming.md b/lectures/phillips_self_confirming.md
index f2a99df..c4a6357 100644
--- a/lectures/phillips_self_confirming.md
+++ b/lectures/phillips_self_confirming.md
@@ -36,6 +36,9 @@ translation:
# 自我确认均衡
+```{index} single: Phillips Curve; Self-Confirming Equilibria
+```
+
```{contents} Contents
:depth: 2
```
@@ -50,9 +53,9 @@ translation:
## 概述
-本讲座完成了 {doc}`phillips_credibility` 中开始的关于菲利普斯曲线权衡的研究。
+本讲座完成了 {doc}`phillips_credibility` 中开始研究的*均衡*部分。
-它遵循 {cite}`Sargent1999` 的第 7 章。
+它遵循 {cite}`Sargent1999` 的第 7 章,在此之后,本系列将从固定信念转向实时学习的信念。
我们寻求这样一些模型:它们与 {doc}`phillips_credibility` 中 {cite}`KydlandPrescott1977` 的基本模型偏离最小,但同时也让政府的*信念*由其自身政策所产生的数据来塑造。
@@ -83,13 +86,13 @@ from scipy.optimize import minimize_scalar
政府相信一个简约形式的菲利普斯曲线,它可以按两个方向拟合:
$$
-\text{古典:} \quad U_t = \gamma' X_{C,t} + \varepsilon_{C,t},
-\qquad X_{C,t} = \begin{bmatrix} y_t & X_{t-1}' \end{bmatrix}',
+\text{古典:} \quad U_t = \gamma^\top X_{C,t} + \varepsilon_{C,t},
+\qquad X_{C,t} = \begin{bmatrix} y_t & X_{t-1}^\top \end{bmatrix}^\top,
$$
$$
-\text{凯恩斯:} \quad y_t = \beta' X_{K,t} + \varepsilon_{K,t},
-\qquad X_{K,t} = \begin{bmatrix} U_t & X_{t-1}' \end{bmatrix}' .
+\text{凯恩斯:} \quad y_t = \beta^\top X_{K,t} + \varepsilon_{K,t},
+\qquad X_{K,t} = \begin{bmatrix} U_t & X_{t-1}^\top \end{bmatrix}^\top.
$$
求解费尔普斯问题以政府的信念 $\gamma$ 为给定,得出通货膨胀的决策规则 $h(\gamma)$。
@@ -112,7 +115,7 @@ $$
U_t = U^* - \frac{\theta}{1 - \rho_2 L}(y_t - x_t) + \frac{v_{1t}}{1 - \rho_1 L},
```
-其中 $|\rho_1| < 1$,$|\rho_2| < 1$,且 $v_t = (v_{1t}, v_{2t})'$ 为向量白噪声,其中 $v_{2t} \equiv y_t - x_t$ 是通货膨胀的意外部分。
+其中 $|\rho_1| < 1$,$|\rho_2| < 1$,且 $v_t = (v_{1t}, v_{2t})^\top$ 为向量白噪声,其中 $v_{2t} \equiv y_t - x_t$ 是通货膨胀的意外部分。
在本讲座的大部分内容中,我们设 $\rho_1 = \rho_2 = 0$ 以阐明理论要点,这将 {eq}`sc_actual` 化简为
@@ -136,7 +139,7 @@ $$
(c)失业率由实际菲利普斯曲线 {eq}`sc_actual` 生成;以及
(d)政府的信念满足最小二乘正交性条件
-$E\left[U_t - \gamma' X_{C,t}\right] X_{C,t}' = 0$(**古典**拟合方向)。
+$\mathbb{E}\left[U_t - \gamma^\top X_{C,t}\right] X_{C,t}^\top = 0$(**古典**拟合方向)。
```
条件(d)使政府的信念依赖于矩矩阵,而这些矩矩阵通过(a)-(c)本身又依赖于政府的信念。
@@ -145,12 +148,18 @@ $E\left[U_t - \gamma' X_{C,t}\right] X_{C,t}' = 0$(**古典**拟合方向)
将(d)替换为**凯恩斯**拟合方向,会得到一个不同的自我确认均衡:
-> (d′)政府拟合凯恩斯菲利普斯曲线,$E\left[y_t - \beta' X_{K,t}\right] X_{K,t}' = 0$,然后通过反演公式 {eq}`sc_invert` 恢复 $\gamma$。
+> (d′)政府拟合凯恩斯菲利普斯曲线,$\mathbb{E}\left[y_t - \beta^\top X_{K,t}\right] X_{K,t}^\top = 0$,然后通过反演公式 {eq}`sc_invert` 恢复 $\gamma$。
由于政府的信念会影响数据的整个概率分布,最小化的方向会影响结果。
```{note}
-一般而言,计算自我确认均衡意味着寻找映射 $\gamma = T(h(\gamma))$(古典)或 $\beta = S(h(\gamma(\beta)))$(凯恩斯)的不动点。正交性条件中的矩是通过求解一个离散李雅普诺夫方程,从该系统的状态空间表示中得到的。实践中,人们迭代一个松弛算法 $\beta_{j+1} = \kappa\beta_j + (1-\kappa) S(\beta_j)$,这与 {doc}`phillips_credibility` 中的最小二乘学习递归相似。
+一般而言,计算自我确认均衡意味着寻找映射 $\gamma = T(h(\gamma))$
+(古典)或 $\beta = S(h(\gamma(\beta)))$(凯恩斯)的不动点。
+
+正交性条件中的矩是通过求解一个离散李雅普诺夫方程,从该系统的状态空间表示中得到的。
+
+实践中,人们迭代一个松弛算法 $\beta_{j+1} = \kappa\beta_j + (1-\kappa) S(\beta_j)$,
+这与 {doc}`phillips_credibility` 中的最小二乘学习递归相似。
```
## 手工可解的特殊情形
@@ -162,9 +171,9 @@ $E\left[U_t - \gamma' X_{C,t}\right] X_{C,t}' = 0$(**古典**拟合方向)
```{math}
:label: sc_moments
-\operatorname{var}(U_t) = \theta^2 \sigma_2^2 + \sigma_1^2,
+\mathbb{V}[U_t] = \theta^2 \sigma_2^2 + \sigma_1^2,
\qquad
-\operatorname{var}(y_t) = \sigma_2^2,
+\mathbb{V}[y_t] = \sigma_2^2,
\qquad
\operatorname{cov}(U_t, y_t) = -\theta \sigma_2^2 .
```
@@ -172,7 +181,7 @@ $E\left[U_t - \gamma' X_{C,t}\right] X_{C,t}' = 0$(**古典**拟合方向)
**古典拟合方向**($U$ 对 $y$):斜率为
$$
-\gamma_1 = \frac{\operatorname{cov}(U_t, y_t)}{\operatorname{var}(y_t)} = -\theta,
+\gamma_1 = \frac{\operatorname{cov}(U_t, y_t)}{\mathbb{V}[y_t]} = -\theta,
$$
而均值必须落在回归线上这一要求给出截距 $\gamma_{-1} = (\gamma_1^2 + 1) U^*$。
@@ -180,7 +189,7 @@ $$
**凯恩斯拟合方向**($y$ 对 $U$):斜率为
$$
-\beta_1 = \frac{\operatorname{cov}(U_t, y_t)}{\operatorname{var}(U_t)}
+\beta_1 = \frac{\operatorname{cov}(U_t, y_t)}{\mathbb{V}[U_t]}
= \frac{-\theta \sigma_2^2}{\sigma_1^2 + \theta^2 \sigma_2^2},
$$
@@ -236,13 +245,19 @@ print(f" γ_1 = {γ1_K:.1f}, γ_(-1) = {γ0_K:.1f}, mean inflation = {y_K:.1f
让我们绘制这两条自我确认的菲利普斯曲线,重现图 7.1。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: The two self-confirming Phillips curves, one for each direction of fit
+ name: fig-sce-two-curves
+---
fig, ax = plt.subplots(figsize=(7, 6))
U_grid = np.linspace(0, 12, 100)
# perceived Phillips curves U = γ_{-1} + γ_1 y => y = (U - γ_{-1}) / γ_1
-ax.plot(U_grid, (U_grid - γ0_C) / γ1_C, 'C0', label='P: classical fit')
-ax.plot(U_grid, (U_grid - γ0_K) / γ1_K, 'C1', label='Q: Keynesian fit')
+ax.plot(U_grid, (U_grid - γ0_C) / γ1_C, 'C0', label='P: classical fit', lw=2)
+ax.plot(U_grid, (U_grid - γ0_K) / γ1_K, 'C1', label='Q: Keynesian fit', lw=2)
ax.plot(sce.U_star, y_C, 'C0o')
ax.annotate('Nash', (sce.U_star, y_C), (sce.U_star + 0.4, y_C - 0.6))
@@ -292,7 +307,7 @@ x_t = C y_{t-1} + (1 - C) x_{t-1}, \quad C \in (0, 1),
其中公众持有常数增益的适应性预期,参数 $C$ 由其*根据数据进行调整*。
-将 $x_t$ 视为状态变量,政府求解费尔普斯问题:通过选择一个反馈规则 $y_t = f_1 + f_2 x_t + v_{2t}$,最大化 $-E_0 \sum_{t=0}^\infty \delta^t\left[(U^* - \theta(y_t - x_t))^2 + y_t^2\right]$。
+将 $x_t$ 视为状态变量,政府求解费尔普斯问题:通过选择一个反馈规则 $y_t = f_1 + f_2 x_t + v_{2t}$,最大化 $-\mathbb{E}_0 \sum_{t=0}^\infty \delta^t\left[(U^* - \theta(y_t - x_t))^2 + y_t^2\right]$。
这恰好是 {doc}`phillips_adaptive` 中带有适应参数 $\lambda = 1 - C$ 的 LQ 费尔普斯问题。
@@ -400,6 +415,12 @@ print(f"Nash inflation θ U* = {mp.θ * mp.U_star:.3f}")
随着政府变得更有耐心,均衡平均通货膨胀率会降至拉姆齐值零附近。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 随着政府变得更有耐心,平均通货膨胀率与均衡增益的变化
+ name: fig-sce-patience
+---
δ_grid = np.array([0.95, 0.96, 0.97, 0.98, 0.99, 0.995])
C_vals, ν_vals = [], []
for δ in δ_grid:
@@ -414,7 +435,7 @@ axes[0].set_xlabel(r'discount factor $\delta$')
axes[0].set_ylabel('mean inflation')
axes[0].legend()
-axes[1].plot(δ_grid, C_vals, 'o-', color='C1')
+axes[1].plot(δ_grid, C_vals, 'o-', color='C1', lw=2)
axes[1].set_xlabel(r'discount factor $\delta$')
axes[1].set_ylabel('equilibrium gain $C$')
@@ -425,7 +446,10 @@ plt.show()
在每个贴现因子下,平均通货膨胀都远低于纳什值,并随着 $\delta \to 1$ 趋近于拉姆齐值零。
```{note}
-精确的均衡值取决于用来保证被感知模型谱密度良好定义的近单位根近似 $\rho$,正如 {doc}`phillips_misspecified` 中所讨论的那样。定性的结论——即结果优于纳什,并随着 $\delta \to 1$ 趋近拉姆齐——是稳健的。
+精确的均衡值取决于用来保证被感知模型谱密度良好定义的近单位根近似 $\rho$,正如
+{doc}`phillips_misspecified` 中所讨论的那样。
+
+定性的结论——即结果优于纳什,并随着 $\delta \to 1$ 趋近拉姆齐——是稳健的。
```
### 谱与脉冲响应
@@ -433,6 +457,12 @@ plt.show()
让我们比较均衡处真实与近似的通货膨胀过程,如 {cite}`Sargent1999` 的图 7.2 和图 7.3 所示。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 均衡处真实与近似通货膨胀过程的谱密度
+ name: fig-sce-spectra
+---
ν_star, F, _ = mp.true_process(C_star)
c_star = mp.best_estimate(C_star)
H = np.abs((1 - (1 - c_star) * mp.z) / (1 - mp.ρ * mp.z))**2
@@ -441,8 +471,8 @@ G = H * σ_ε2
half = mp.N // 2
fig, ax = plt.subplots(figsize=(8, 5))
-ax.plot(mp.ω[:half], np.log(F[:half]), 'C0', label='true model')
-ax.plot(mp.ω[:half], np.log(G[:half]), 'C1--', label='approximating model')
+ax.plot(mp.ω[:half], np.log(F[:half]), 'C0', label='true model', lw=2)
+ax.plot(mp.ω[:half], np.log(G[:half]), 'C1--', label='approximating model', lw=2)
ax.set_xlabel(r'angular frequency $\omega$')
ax.set_ylabel('log spectral density')
ax.legend()
@@ -454,6 +484,12 @@ plt.show()
真实的通货膨胀率仅具有中等程度的序列相关性,而——正如 {doc}`phillips_misspecified` 中的布雷模型一样——近似模型使用单位根来模拟均值,即用二阶矩来捕捉一阶矩。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: 真实与近似通货膨胀过程的脉冲响应
+ name: fig-sce-irf
+---
def ima_impulse(num, den, T=25):
"IRF of (1 - num L)/(1 - den L)."
h = np.empty(T)
@@ -468,8 +504,8 @@ irf_true = ima_impulse(1 - C_star, ψ)
irf_approx = ima_impulse(1 - c_star, mp.ρ)
fig, ax = plt.subplots(figsize=(8, 5))
-ax.plot(irf_true, 'C0o-', ms=4, label='true model')
-ax.plot(irf_approx, 'C1s--', ms=4, label='approximating model')
+ax.plot(irf_true, 'C0o-', ms=4, label='true model', lw=2)
+ax.plot(irf_approx, 'C1s--', ms=4, label='approximating model', lw=2)
ax.set_xlabel('lag')
ax.set_ylabel('response')
ax.legend()
@@ -521,6 +557,7 @@ ax.plot(σ1_grid, y_keynes, label='Keynesian mean inflation')
ax.axhline(5.0, color='k', ls='--', lw=1, label='Nash')
ax.set_xlabel(r'$\sigma_1$')
ax.set_ylabel('mean inflation')
+ax.set_title('Keynesian mean inflation by shock size')
ax.legend()
plt.show()
```
@@ -560,6 +597,7 @@ ax.plot(C_star, C_star, 'ko')
ax.annotate('equilibrium', (C_star, C_star), (C_star + 0.03, C_star - 0.03))
ax.set_xlabel('$C$')
ax.set_ylabel('$B(C)$')
+ax.set_title('The best-estimate map and the equilibrium gain')
ax.legend()
plt.show()
```
@@ -567,4 +605,4 @@ plt.show()
最优估计映射在均衡增益处与 45 度线相交,证实了 $C = B(C)$。
```{solution-end}
-```
\ No newline at end of file
+```
diff --git a/lectures/phillips_two_stories.md b/lectures/phillips_two_stories.md
index 67f904d..18f57fb 100644
--- a/lectures/phillips_two_stories.md
+++ b/lectures/phillips_two_stories.md
@@ -13,6 +13,7 @@ translation:
title: 美国通货膨胀的兴衰
headings:
Overview: 概述
+ Overview::A note on notation: 关于记号的说明
Facts: 事实
The Phillips curve in the data: 数据中的菲利普斯曲线
Two interpretations: 两种解释
@@ -45,6 +46,9 @@ translation:
# 美国通货膨胀的兴衰
+```{index} single: Phillips Curve; Rise and Fall of U.S. Inflation
+```
+
```{contents} Contents
:depth: 2
```
@@ -74,13 +78,15 @@ translation:
两个故事的区别在于该理论是如何被采纳的:
* **自然率理论的胜利。** 学术经济学家发现了自然率假说,指出任何通货膨胀与失业之间的权衡都是暂时的,并最终说服政策制定者追求低通货膨胀。
-* **计量经济学政策评估的平反。** 政策制定者从未放弃罗伯特·卢卡斯在其著名批判中所抨击的方法。他们反复重新估计菲利普斯曲线并用它来选择目标,而正是*数据本身*——一条不断向不利方向漂移的经验菲利普斯曲线——引导他们走向更低的通货膨胀。
+* **计量经济学政策评估的平反。** 政策制定者从未放弃罗伯特·卢卡斯在其著名批判中所抨击的方法。
+ - 他们反复重新估计菲利普斯曲线并用它来选择目标,正是*数据本身*——一条不断向不利方向漂移的经验菲利普斯曲线——引导他们走向更低的通货膨胀。
本讲座介绍了支持这两个故事的事实,勾勒出这两种解释,并回顾了第二章既援引又修正的卢卡斯批判。
该系列的其余讲座建立了相应的模型:
* {doc}`phillips_credibility` —— 单期基德兰德-普雷斯科特(Kydland-Prescott)可信度问题(第三章)。
+* {doc}`phillips_credible_policies` —— 重复经济中的声誉问题,以及为何可信政策理论用不可知论取代了悲观主义(第四章)。
* {doc}`phillips_adaptive` —— 适应性预期与菲尔普斯问题(第五章)。
* {doc}`phillips_misspecified` —— 最优错误设定信念下的均衡(第六章)。
* {doc}`phillips_self_confirming` —— 自我确认均衡(第七章)。
@@ -90,6 +96,39 @@ translation:
* {doc}`phillips_lost_conquest` —— 将同样的工具应用于 2020 年代的通货膨胀以及美联储的迟缓反应({cite}`SargentWilliams2025`)。
* {doc}`phillips_drifts_volatilities` —— 一篇实证后记,将带漂移系数、随机波动率的向量自回归模型拟合到数据上,探讨大通胀究竟是政策不当还是运气不佳造成的({cite}`CogleySargent2005`)。
+(phillips_notation)=
+### 关于记号的说明
+
+各讲座遵循其各自所依据来源的记号约定,而这些来源彼此之间并不一致。
+
+与其强加一套统一的记号方案、从而使读者难以对照可能想要参阅的原始论文,我们选择让每一讲都忠实于其来源,并在此记录各讲之间的记号对应关系。
+
+| 对象 | 符号 | 出处 |
+|---|---|---|
+| 通货膨胀 | $y$ | {doc}`phillips_credibility` 至 {doc}`phillips_self_confirming` |
+| | $\pi$ | 从 {doc}`phillips_escaping_nash` 起 |
+| 公众的预期通货膨胀 | $x$ | {doc}`phillips_credibility`、{doc}`phillips_adaptive` |
+| 政府的系统性通货膨胀 | $x$ | {doc}`phillips_escaping_nash`、{doc}`phillips_priors` |
+| 自然失业率 | $U^*$ | {doc}`phillips_credibility` 至 {doc}`phillips_learning` |
+| | $u$ | {doc}`phillips_escaping_nash`、{doc}`phillips_priors` |
+| 菲利普斯曲线斜率 | $\theta$ | {doc}`phillips_credibility` 至 {doc}`phillips_escaping_nash` |
+| 政府的信念 | $\gamma$ | {doc}`phillips_adaptive` 至 {doc}`phillips_priors` |
+| | $\theta$ | {doc}`phillips_lost_conquest`、{doc}`phillips_drifts_volatilities` |
+| 贴现因子 | $\delta$ | {doc}`phillips_credible_policies` 至 {doc}`phillips_learning` |
+| | $\beta$ | {doc}`phillips_lost_conquest`、{doc}`phillips_drifts_volatilities` |
+| 学习增益 | $\lambda$、$g_t$ | {doc}`phillips_adaptive`、{doc}`phillips_learning` |
+| | $\varepsilon$ | {doc}`phillips_escaping_nash`、{doc}`phillips_priors` |
+
+有三处符号冲突值得提前指出,因为如果读者把符号含义一路带下去,就会在这些地方被绊住。
+
+字母 $x$ 换了阵营:在 {doc}`phillips_credibility` 中它表示*公众*的预期,而从 {doc}`phillips_escaping_nash` 起,它表示的是*政府*所设定的量。
+
+二者在均衡中恰好相等,这正是这一转变容易被忽略的原因。
+
+字母 $\theta$ 在前几讲中是菲利普斯曲线的斜率,而在最后两讲中则是政府的整个信念向量。
+
+而 $\lambda$ 则身兼四职:在 {doc}`phillips_adaptive` 中是凯根-弗里德曼适应性参数,在 {doc}`phillips_learning` 中是遗忘因子,在 {doc}`phillips_priors` 中是先验紧缩参数,在 {doc}`phillips_lost_conquest` 中则是实测的持续性根。
+
让我们从一些导入开始:
```{code-cell} ipython3
@@ -103,7 +142,9 @@ from statsmodels.tsa.filters.bk_filter import bkfilter
```{note}
接下来两节中的图表复现了 {cite}`Sargent1999` 第一章和第二章中的图表,这些图表所用的数据取自 20 世纪 90 年代末。
+
我们从 [FRED](https://fred.stlouisfed.org/) 下载相应的原始数据序列,并将关注范围限定在与原书相同的历史区间内。
+
之后的 {ref}`phillips_after_1999` 一节会将其中最具启发性的图表延伸到当下,并探讨这额外的四分之一个世纪的数据对这两个故事意味着什么。
```
@@ -124,6 +165,12 @@ inflation_ma = inflation.rolling(13, center=True).mean()
```
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: Monthly CPI inflation, 13-month centered moving average, 1948-1999
+ name: fig-ts-inflation
+---
fig, ax = plt.subplots(figsize=(9, 5))
ax.plot(inflation_ma, lw=1.2)
ax.axhline(0, color='k', lw=0.5)
@@ -159,6 +206,12 @@ data.head()
图 1.2 将这两个原始序列绘制在一起。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: Monthly unemployment (white men 20+) and inflation rates
+ name: fig-ts-raw-series
+---
fig, ax = plt.subplots(figsize=(9, 5))
ax.plot(data.index, data['inflation'], 'C0', lw=1, label='inflation (CPI)')
ax.plot(data.index, data['unemployment'], 'C1:', lw=1.2,
@@ -182,6 +235,12 @@ bk.columns = ['inflation_cycle', 'unemployment_cycle']
```
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: Business-cycle components of inflation and unemployment, Baxter-King bandpass filter
+ name: fig-ts-bandpass
+---
fig, ax = plt.subplots(figsize=(9, 5))
ax.plot(bk.index, bk['inflation_cycle'], 'C0', lw=1, label='inflation')
ax.plot(bk.index, bk['unemployment_cycle'], 'C1:', lw=1.2,
@@ -202,6 +261,12 @@ plt.show()
图 1.4 将原始序列相互对照绘制,图 1.5 则展示了商业周期成分。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "Inflation against unemployment, 1960-1982: raw series and business-cycle components"
+ name: fig-ts-scatter-6082
+---
sub = slice('1960', '1982')
fig, axes = plt.subplots(1, 2, figsize=(12, 5))
@@ -227,7 +292,11 @@ plt.show()
图 1.5 揭示了**菲利普斯回路**:通货膨胀与失业率描绘出的是逆时针方向的回路,而非单一的稳定曲线,这正是自然率理论所强调的、位于故事核心的预期转变的特征表现。
```{note}
-原书通过选取单一失业率序列来调整人口结构变化的影响。若采用更广义的失业定义,就会引入额外的低频人口结构成分,人们可能会用单位根过程来对其建模。而本文则从另一个来源——即摆脱布雷顿森林体系约束后货币当局*漂移的信念*——将单位根引入通货膨胀-失业率过程之中。
+原书通过选取单一失业率序列来调整人口结构变化的影响。
+
+若采用更广义的失业定义,就会引入额外的低频人口结构成分,人们可能会用单位根过程来对其建模。
+
+而本文则从另一个来源——即摆脱布雷顿森林体系约束后货币当局*漂移的信念*——将单位根引入通货膨胀-失业率过程之中。
```
## 两种解释
@@ -307,7 +376,7 @@ plt.show()
尽管政府的不变性假设是错误的,但它并未在结果中受挫,因为这些结果在统计上与它的信念是一致的。
-自我确认均衡是一种理性预期均衡,但其自由参数比卢卡斯所使用的模型*更少*——而恰恰正是这些缺失的参数,才是表现制度变迁所需要的。
+在自我确认均衡中,政府的信念*沿均衡路径*是正确的,因此其预测满足与理性预期均衡相同的跨方程约束;但相对于卢卡斯所使用的完全结构化模型而言,政府的模型所包含的自由参数*更少*——而恰恰正是这些缺失的参数,才是表现制度变迁所需要的。
要容纳制度变迁和漂移系数,就必须*抵制*向自我确认均衡的收敛。
@@ -339,6 +408,8 @@ plt.show()
单期的 {cite}`KydlandPrescott1977` 模型得出了一个悲观的预测——即高通货膨胀的时间一致(纳什)结果——但可信政策理论的重复经济版本,用*不可知论*取代了这种悲观:太多的结果都变得可维持,以至于该理论只能给出微弱的预测。
+{doc}`phillips_credible_policies` 证明了这一点,它用 {cite}`APS1990` 的递归方法计算出可维持数值的整个集合,并展示了三种带来相同收益却截然不同的均衡。
+
这种弱预测性是我们在宣称自然率理论取得胜利之前应当迟疑的第一个理由。
随后我们从卢卡斯批判处折返,重新从费尔普斯基准出发,但作一处改动:政府对私人部门的模型不再是任意的——它是*拟合历史数据*得到的。
@@ -351,7 +422,7 @@ plt.show()
这些适应性模型是对理性预期一种*有节制的*退让,而非对其的彻底抛弃。
-它们不包含控制预期的自由参数;在每一期,它们都施加与理性预期模型相同的跨方程约束;并且——由于自我确认均衡是其*均值动态*的吸引子——它们在平静的条件下会收敛回理性预期,满足了 {cite}`Kreps1998` 所提出的一个诉求。
+它们不包含控制预期的自由参数;在每一期,它们都施加与理性预期模型相同的跨方程约束;并且——由于自我确认均衡是其*均值动态*的吸引子——在平静的条件下它们会收敛回自我确认均衡,从而在均衡路径上收敛回理性预期,满足了 {cite}`Kreps1998` 所提出的一个诉求。
但是,遵循 {cite}`Sims1988` 的思路,我们真正感兴趣的是适应性所带来的*周期性*动态。
@@ -427,6 +498,12 @@ inflation_yoy = 100 * (cpi_full / cpi_full.shift(12) - 1)
将其延伸至今,增添了原书所无法看到的三段历程:从 20 世纪 80 年代中期开始、通货膨胀低而稳定的*大缓和*时期;2008 年金融危机后,一段长期接近于零、并一度略低于零的时期;以及 2021-2022 年一次骤然飙升至 1981 年以来最高水平、随后又迅速回落的过程。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: Inflation extended to the present, with the book's window shaded
+ name: fig-ts-inflation-long
+---
fig, ax = plt.subplots(figsize=(11, 5))
ax.plot(inflation_ma_full, lw=1)
ax.axhline(0, color='k', lw=0.5)
@@ -457,6 +534,12 @@ plt.show()
将其延伸后,展示了新数据中最引人注目的两大宏观经济事件:2020 年新冠疫情导致的失业率骤升——一度是大萧条以来的最高水平——以及随之而来的通货膨胀飙升。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: Unemployment and inflation since 1990
+ name: fig-ts-recent
+---
recent = slice('1990', None)
fig, ax = plt.subplots(figsize=(11, 5))
@@ -485,6 +568,12 @@ plt.show()
我们将样本划分为原书所涉及的加速时期、大缓和时期,以及 2008 年之后的时期,并在每个时期分别绘制通货膨胀对失业率的散点图。
```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: The inflation-unemployment scatter across three eras
+ name: fig-ts-three-eras
+---
scatter_data = pd.concat([inflation_yoy.rename('inflation'),
u_full.rename('unemployment')], axis=1).dropna()
@@ -504,8 +593,6 @@ plt.tight_layout()
plt.show()
```
-这三片散点云的样貌几乎不可能相差更大。
-
在 1960-1983 年间,数据点散布在一个很宽的通货膨胀率范围内——这正是预期不断变化、菲利普斯曲线呈现*回路*形态的时代。
在 1984-2007 年间,它们收缩成一团紧密、低位、近乎平坦的点云——这就是大缓和时期,此时通货膨胀几乎对失业率毫无反应。
@@ -531,12 +618,16 @@ plt.show()
2020 年之前那十年近乎零通货膨胀的时期——那条看似*平坦*的菲利普斯曲线,既没有出现 2009-2013 年"消失的反通货膨胀",也没有出现 2015-2019 年"消失的通货膨胀",都不符合一条稳定曲线——恰恰正是本书中适应性政府会实时追踪其斜率和截距不断变化的那种漂移型经验关系。
```{note}
-本书自身也会坚持提出这样一个警示:其机制假设*基本面*——即真实的数据生成过程——是稳定的,因此所有的作用都来自政府不断演化的信念。而 2021-2022 年这一事件涉及了真实的供给冲击(疫情引发的中断、能源价格),这已超出了该假设的范围。将信念的转变与基本面的转变区分开来,正是使这段历史如此难以捉摸、也如此引人入胜的那个识别问题。
+本书自身也会坚持提出这样一个警示:其机制假设*基本面*——即真实的数据生成过程——是稳定的,因此所有的作用都来自政府不断演化的信念。
+
+2021-2022 年这一事件涉及了真实的供给冲击(疫情引发的中断、能源价格),这已超出了该假设的范围。
+
+将信念的转变与基本面的转变区分开来,正是使这段历史如此难以捉摸、也如此引人入胜的那个识别问题。
```
本系列后续讲座中所建立的工具——自我确认均衡、漂移系数以及逃逸动态——仍然是探讨新数据所提出的这一问题的一种自然语言:一个可信的低通货膨胀均衡,是否会在每次冲击之后重新锚定,还是说一连串的意外仍可能使信念重新开始漂移,就像 1965 年之后所发生的那样?
-最后一讲,{doc}`phillips_lost_conquest`,正是把这些工具应用于 2021-2022 年的飙升,并探讨了为什么美联储的反应如此迟缓。
+{doc}`phillips_lost_conquest` 正是把这些工具应用于 2021-2022 年的飙升,并探讨了为什么美联储的反应如此迟缓,而收尾一讲 {doc}`phillips_drifts_volatilities` 则直接从数据出发,探讨战后宏观经济动态的系数究竟是否真的发生了漂移。
## 练习
@@ -573,4 +664,4 @@ print(f"business-cycle correlation, 1960-82: {corr_cycle:+.2f}")
一旦滤除这些低频移动,一种强烈的负相关关系便显现出来,证实了菲利普斯权衡关系是在商业周期频率上发挥作用的。
```{solution-end}
-```
\ No newline at end of file
+```
diff --git a/lectures/prospects_bounded_rationality.md b/lectures/prospects_bounded_rationality.md
new file mode 100644
index 0000000..7a2700d
--- /dev/null
+++ b/lectures/prospects_bounded_rationality.md
@@ -0,0 +1,559 @@
+---
+jupytext:
+ text_representation:
+ extension: .md
+ format_name: myst
+ format_version: 0.13
+ jupytext_version: 1.17.1
+kernelspec:
+ display_name: Python 3 (ipykernel)
+ language: python
+ name: python3
+translation:
+ title: 1993 年对宏观经济学中有限理性前景的展望
+ headings:
+ Overview: 概览
+ The quest, restated: 追问,重述
+ 'The debit side: how much we must hard-wire': 借方一栏:我们必须预先设定多少内容
+ Why the econometricians have not returned the compliment: 为何计量经济学家未曾回礼
+ Why the econometricians have not returned the compliment::The nuisance parameter, made concrete: 具体化的滋扰参数
+ 'The credit side: selection, computation, and a returned gift': 贷方一栏:选择、计算,以及一份被回赠的礼物
+ Reading the ledger against the quest: 对照追问来解读这份账本
+ 'Postscript: what became of the program': 后记:这一研究计划后来的走向
+ 'Postscript: what became of the program::The arbitrariness was disciplined': 任意性被规训了
+ 'Postscript: what became of the program::The transition dynamics arrived, in a narrower form than hoped': 转型动态理论确实到来了,但形式比人们所期望的要狭窄
+ 'Postscript: what became of the program::The econometricians did return the compliment': 计量经济学家终究回礼了
+ 'Postscript: what became of the program::A second retreat, made differently': 第二次以不同方式进行的撤退
+ 'Postscript: what became of the program::The algorithms kept crossing over': 算法持续跨界流动
+ 'Postscript: what became of the program::Reading the ledger again': 再度解读这份账本
+---
+
+(prospects_bounded_rationality)=
+```{raw} jupyter
+
+```
+
+# 1993 年对宏观经济学中有限理性前景的展望
+
+```{index} single: Bounded Rationality; Prospects
+```
+
+```{contents} Contents
+:depth: 2
+```
+
+## 概览
+
+本讲是本系列的收官讲座,其性质与其他各讲不同。
+
+它不携带任何新模型。
+
+相反,它汇集了 {cite:t}`Sargent1993` 于 1993 年记录下的判断——他在结束《宏观经济学中的有限理性》一书时所留下的观点、保留意见和期望——并将它们与本书开篇、也是本系列第一讲 {doc}`bounded_rationality` 开篇所提出的追问联系起来。
+
+那个追问有一个目标:一套关于**转型动态**的理论,即 1989 年东欧改革者们在没有地图的情况下不得不应对的非均衡调整过程。
+
+其路径是将理性主体从我们的模型中驱逐出去,代之以行为像**计量经济学家**一样的"人工智能"主体——他们收集数据、形成理论、进行估计并不断适应。
+
+我们现在可以追问:截至 1993 年,这条路径把我们带到了离目标多远的地方。
+
+萨金特自己的答案是一份账本,正反两方均有记录。
+
+在贷方一栏:适应性动态作为在多重均衡中进行**选择**的手段,以及作为**计算**均衡的工具——演化编程。
+
+在借方一栏:最初的奖赏——一套转型动态理论——在很大程度上仍未兑现;而且出现了一个引人注目的不对称现象。
+
+这一研究计划的初衷是让我们模型中的主体表现得更像计量经济学家。
+
+萨金特观察到,计量经济学家们却并未以同样的方式回礼。
+
+本讲将解释这一不对称现象,赋予它一个具体的数值形象,并对照最初的追问来解读这份账本。
+
+本讲最后附有一篇后记,讲述该研究计划在 1993 年之后的演变。
+
+让我们从一些导入开始。
+
+```{code-cell} ipython3
+import numpy as np
+import matplotlib.pyplot as plt
+```
+
+## 追问,重述
+
+有必要重述第一讲的论证,因为下文的一切都是以此为标准来衡量的。
+
+理性预期提出了两项要求:个体理性,以及认知的相互一致性。
+
+第二项要求是苛刻的一项——这正是问题的关键——当一个理性预期模型被拿去与数据对照时,它赋予模型内部主体的知识,远超过研究他们的计量经济学家所拥有的知识。
+
+主体们是利用*均衡*概率分布来求解他们的欧拉方程的,而这恰恰是计量经济学家仍在苦苦估计的那些分布。
+
+有限理性研究计划试图通过将主体降至计量经济学家自身的水平来弥合这一差距:他们也必须一边前进一边从数据中学习这些分布。
+
+人们曾希望这能带来理性预期无法提供的东西:对系统*仍处于调整过程中*的描述,即在信念与结果尚未达成相互一致之前的状态。
+
+那正是一套转型动态理论应有的样子。
+
+我们现在依照 {cite:t}`Sargent1993` 自己的记账方式来盘点:先看借方,再看核心的不对称现象,最后看贷方。
+
+## 借方一栏:我们必须预先设定多少内容
+
+第一项保留意见关乎**任意性**。
+
+有限理性最容易通过其*反面*——理性预期——来定义,而这种可塑性本身就是一个负担。
+
+一旦我们不再坚持主体知晓均衡,我们就必须逐案决定:他们究竟*知道*什么,又是如何学习的。
+
+他们是知道自己的效用函数和利润函数,还是也必须学习这些?
+
+他们是懂微积分和动态规划,还是仅凭试错?
+
+他们是仅从自身经验中学习,还是也从他人经验中学习?
+
+我们把他们所学限定在哪一类近似函数之中?
+
+本系列中的每一个模型都是通过**预先设定**来回答这些问题的,即对主体进行大量提示,并盯着我们所期望其达到的结果。
+
+布雷模型中的主体——在 {doc}`bounded_rationality` 的适应性预期模型以及最小二乘学习文献中——知道正确的供给曲线,只需估计一个条件期望值代入其中即可。
+
+马塞特-萨金特模型中的主体懂得动态规划,知道自己的回报函数,也知道运动法则的参数形式——他们仅仅缺少其系数,而系数是通过向量自回归来更新的。
+
+即便是 {doc}`marimon_mcgrattan_sargent` 中受到提示要少得多的分类器主体——他们从未被告知自己的效用函数,只有在体验到效用时才能识别它——仍然被告知*何时*做出选择、*以何种*信息为条件,而其记账系统与遗传算子的整个装置都是手工设计的,并以基约塔基-赖特均衡为目标。
+
+第二项保留意见关乎**简单性**。
+
+我们赋予主体的学习任务,与一门初级计量经济学课程相比都显得微不足道,更不用说现实中的企业和家庭实际上正在隐含求解的那些任务了。
+
+我们要求一个主体学习单一的、时不变的决策规则,或一组固定的条件期望。
+
+我们并未要求它去学习一个联立方程系统的参数,或去推断从政策体制到分布的映射——而这些正是计量经济学困难所在。
+
+综合来看,这些保留意见直接触及最初的奖赏。
+
+我们让适应性主体置身其中的环境,比我们真正关心的那些转型过程要稳定和友善得多。
+
+收敛速度的结果十分稀少,而可处理性又迫使我们对主体信念的分布施加严格的限制。
+
+因此,萨金特在 1993 年判断,适应性过程的文献远未能为一套实时转型动态理论——正是这场追问所要寻找的东西——提供牢靠的基础。
+
+不过,他并不愿以这一失败作为结语,他所选用的棒球比喻是刻意谦逊的:
+
+> 若以适应性方法迄今未能"打出全垒打"、即未能给出一套良好的转型动态理论这一失败作为本文的结尾,那既不明智,也不公平。转型动态问题由来已久、困难重重。因此,或许可以将这些方法加深了我们对该问题的理解,视为打出了一垒安打,或者至少是一记高飞牺牲打。
+
+## 为何计量经济学家未曾回礼
+
+这便是赋予本讲以主题的不对称现象。
+
+有限理性研究计划归根结底是一场运动,旨在让我们模型中的主体表现得更像那些构建并估计这些模型的计量经济学家。
+
+模仿是最真诚的恭维。
+
+因此我们本可以预期,宏观计量经济学家会争相将这些模型拟合到数据上。
+
+然而并未出现这样的争相追捧。
+
+{cite:t}`Chung1990` 对西姆斯政策制定者模型——即 {doc}`olg_adaptive_money` 中的应用——的估计,萨金特指出,几乎是他所知道的*唯一*一个在计量经济学上认真对待有限理性的宏观经济应用。
+
+为何如此冷淡?
+
+其中的原因值得详细说明,因为这并非品味问题。
+
+应用计量经济学家所遵循的至理名言是卢卡斯的告诫:*要警惕携带自由参数而来的理论家*。
+
+用一个有限理性主体替代一个理性主体,会**增加**参数:描述信念以及信念如何变动的参数。
+
+以最简单的情形——布雷模型为例。
+
+相对于其理性预期版本,适应性版本至少增加了初始信念,以及一个设定增益序列的参数,而人们可能还想用更多参数来描述增益的形态。
+
+这已经足以让卢卡斯的警告有的放矢。
+
+但还有一个更深层的问题,这也正是为何这些新增参数不仅不受欢迎,而且实际上难以估计的原因。
+
+由于适应性系统会**收敛**到理性预期均衡,这些额外的参数仅仅影响*瞬态过程*。
+
+数据的渐近分布不包含关于它们的任何信息。
+
+让我们直接看一看。
+
+### 具体化的滋扰参数
+
+布雷的蛛网经济 {cite:p}`Bray1982` 通过下式设定市场价格:
+
+$$
+p_t = a + b\, \beta_t + u_t,
+$$
+
+其中 $\beta_t$ 是主体预期的价格,通过对过去价格取平均而形成:
+
+$$
+\beta_t = \beta_{t-1} + \gamma_t\,(p_{t-1} - \beta_{t-1}),
+\qquad
+\gamma_t = \frac{1}{t + t_0},
+$$
+
+$u_t$ 是一个独立同分布的冲击。
+
+常数 $t_0$ 是主体赋予其初始信念的权重,以观测值个数计量:当 $t_0 = 50$ 时,他们将 $\beta_0$ 视为总结了此前五十个价格,因此经过 $t$ 期之后,$\beta_0$ 依然带有 $t_0/(t + t_0)$ 的权重。
+
+这样的权重是必要的,否则这个练习将毫无内容可言。
+
+当 $t_0 = 0$ 时,首个增益为 $\gamma_1 = 1$,因此 $\beta_1 = p_0$ 恰好成立,初始信念在一期之后便被抹去——计量经济学家将没有任何瞬态过程可供尝试估计。
+
+当 $b < 1$ 时,无论起始值为何,信念都会收敛到理性预期值 $\beta^\star = a/(1-b)$。
+
+```{code-cell} ipython3
+a, b, sigma_u = 5.0, 0.7, 1.0
+t0 = 50 # weight on the initial belief
+β_star = a / (1 - b) # rational expectations belief
+
+def simulate(β0, u):
+ "Bray's cobweb under least squares learning, given a shock path u."
+ T = len(u)
+ β = np.empty(T)
+ p = np.empty(T)
+ β[0] = β0
+ for t in range(T):
+ p[t] = a + b * β[t] + u[t]
+ if t + 1 < T:
+ β[t + 1] = β[t] + (1 / (t + 1 + t0)) * (p[t] - β[t])
+ return p, β
+
+print(f"rational expectations belief β* = {β_star:.3f}")
+```
+
+初始信念 $\beta_0$ 就是这个额外的"有限理性"参数。
+
+来看三个初始信念迥异的经济体如何遗忘各自的起点。
+
+```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "The belief forgets its starting point"
+ name: fig-pbr-belief
+---
+rng = np.random.default_rng(0)
+u = rng.standard_normal(400)
+
+fig, ax = plt.subplots(figsize=(7.5, 4))
+for β0, colour in [(2.0, 'C0'), (16.667, 'C1'), (40.0, 'C2')]:
+ _, β = simulate(β0, u)
+ ax.plot(β, color=colour, lw=1.3, label=fr"$\beta_0 = {β0}$")
+ax.axhline(β_star, color='k', ls='--', lw=0.8, label=r"$\beta^\star$")
+ax.set_xlabel("$t$")
+ax.set_ylabel(r"belief $\beta_t$")
+ax.legend(frameon=False)
+plt.show()
+```
+
+三者都收敛到了 $\beta^\star$。
+
+现在换上计量经济学家的视角。
+
+给定一个价格样本以及结构参数 $(a, b)$,该模型对任何候选初始信念 $\beta_0$ 都蕴含着一个冲击 $\hat u_t = p_t - a - b\,\beta_t(\beta_0)$,因为信念路径由 $\beta_0$ 和观测到的价格共同确定。
+
+隐含冲击的平方和衡量了给定的 $\beta_0$ 与数据的拟合程度。
+
+```{code-cell} ipython3
+def belief_path(β0, p):
+ "Belief sequence implied by an initial belief and an observed price path."
+ T = len(p)
+ β = np.empty(T)
+ β[0] = β0
+ for t in range(1, T):
+ β[t] = β[t - 1] + (1 / (t + t0)) * (p[t - 1] - β[t - 1])
+ return β
+
+def ssr(β0, p):
+ "Sum of squared implied shocks, as a function of the belief parameter β0."
+ β = belief_path(β0, p)
+ return np.sum((p - a - b * β) ** 2)
+```
+
+数据能否将 $\beta_0$ 确定下来,是一个关于*信息*随着样本增大而积累速度的问题。
+
+自然的判读方式是看**超额**残差平方和——一个错误的 $\beta_0$ 相对于最优值拟合得有多差。
+
+对于一个普通的、可良好识别的参数而言,每一个新的观测值都会增加对错误取值的惩罚,因此超额值会随 $T$ 成比例增长,置信区间会以 $1/\sqrt{T}$ 的速度收缩。
+
+来看看这里会发生什么。
+
+```{code-cell} ipython3
+---
+mystnb:
+ figure:
+ caption: "Information about the belief parameter stops accumulating"
+ name: fig-pbr-excess
+---
+grid = np.linspace(2, 40, 80)
+
+def excess_curve(T, seed=1):
+ "SSR(β₀) − min SSR across the β₀ grid, for a sample of length T."
+ u = np.random.default_rng(seed).standard_normal(T)
+ p, _ = simulate(β_star, u)
+ curve = np.array([ssr(b0, p) for b0 in grid])
+ return curve - curve.min()
+
+fig, ax = plt.subplots(figsize=(7.5, 4))
+for T, colour in zip((50, 500, 5000), ('C0', 'C1', 'C2')):
+ ax.plot(grid, excess_curve(T), color=colour, lw=1.5, label=f"$T = {T}$")
+ax.axvline(β_star, color='k', ls='--', lw=0.8)
+ax.set_xlabel(r"belief parameter $\beta_0$")
+ax.set_ylabel("excess sum of squares")
+ax.legend(frameon=False)
+plt.show()
+```
+
+这些曲线彼此重叠,而不是变得越来越陡峭。
+
+为了说明这并非所绘范围造成的假象,我们将错误 $\beta_0$ 所受的惩罚,与同一模型中另一个普通结构参数——错误的*斜率* $b$——从同样数据中所受的惩罚作一比较。
+
+```{code-cell} ipython3
+rows = []
+for T in (50, 500, 5_000, 50_000):
+ u = np.random.default_rng(1).standard_normal(T)
+ p, _ = simulate(β_star, u)
+ β = belief_path(β_star, p) # belief path at the true β0
+ wrong_β0 = ssr(2.0, p) - ssr(β_star, p)
+ wrong_slope = (np.sum((p - a - 0.75 * β) ** 2)
+ - np.sum((p - a - b * β) ** 2))
+ rows.append([T, wrong_β0, wrong_slope])
+
+for T, e_b0, e_b in rows:
+ print(f"T = {T:6d}: penalty for β₀ = 2 : {e_b0:10.1f} "
+ f"penalty for b = 0.75 : {e_b:12.1f}")
+```
+
+这两列的表现完全不同。
+
+斜率取错所受到的惩罚随样本量成比例增长:数据量增加一百倍,就会带来对错误取值强一百倍的反证力度,这正是可识别性应有的样子。
+
+而初始信念取错所受到的惩罚却停止了增长。
+
+超过几千个观测值之后,它就被固定在一个常数上,此后每增加一次观测都无法提供关于 $\beta_0$ 的信息。
+
+$\beta_0$ 的置信区间永远不会收缩;该参数根本无法被一致地估计出来。
+
+这正是问题的技术核心所在。
+
+有限理性所增添的参数完全存在于瞬态过程之中。
+
+无论我们此后观察经济多长时间,一个瞬态过程所贡献的信息量都是固定而有限的,因此这些参数就变成了**估计上的滋扰**——它们进入似然函数,而数据对它们能够说明的内容始终是有界的。
+
+而且还有一个最终的、决定性的原因,可以解释计量经济学家们冷淡反应。
+
+许多应用宏观计量经济学家所寻求的方法,恰恰是要*减少*解释数据所需的参数数目。
+
+而有限理性所提供的恰恰不是减少,而是相反。
+
+它提供的是更多参数,其中大多数还识别得很弱。
+
+于是,这份恭维只是单向而行。
+
+理论家们按照计量经济学家的形象重塑了他们的模型主体;而计量经济学家们面对充满额外、识别薄弱参数的模型,且已被卢卡斯明确告诫要警惕正是此类情况,便婉拒了这份礼物。
+
+## 贷方一栏:选择、计算,以及一份被回赠的礼物
+
+这份账本并非单方面的。
+
+与前述借方相对,萨金特列出了三项真正的成就,其中最后一项悄然抵消了刚刚描述的那种不对称现象的一部分。
+
+**均衡选择。**
+
+在一个理性预期模型存在多重均衡的地方,一个由适应性主体构成的系统往往会收敛到*某一个特定的*均衡,从而使学习成为在多重均衡中进行选择的手段。
+
+我们已多次见证这一点: {doc}`olg_adaptive_money` 中被选中的低通胀均衡、 {doc}`exchange_rate_learning` 中依赖历史路径的汇率、以及 {doc}`marimon_mcgrattan_sargent` 中的基本货币均衡。
+
+萨金特坦率地承认,他对这种用法的偏爱,与他对执行这种选择的动态过程本身的怀疑之间,存在某种张力:
+
+> 我知道,对实时动态过程心存怀疑、却又保留由其所选中的均衡,这在逻辑上是不一致的。我承认,我对第六章所描述的货币模型中所执行的那种选择所抱有的偏爱,部分源于我先验地相信被选中的那些均衡在我看来是合理的。
+
+**演化编程。**
+
+如果一个由适应性主体构成的群体能够可靠地收敛到某个均衡,我们就可以将该群体作为一种*计算*该均衡的方法来运行,尤其是在那些复杂到无法手工求解的模型中。
+
+{doc}`marimon_mcgrattan_sargent` 正是这样处理五种商品的基约塔基-赖特经济的,因为当时并没有现成的解析刻画方法。
+
+萨金特预期这种用法将会被频繁采用。
+
+**为计量经济学家提供的新工具。**
+
+在这里,不对称现象出现了反转。
+
+关于并行算法和遗传算法的文献,为计量经济学家提供了新的计算工具——其中包括遗传算法和随机高斯-牛顿程序——用以求解他们*自身*的估计和优化问题。
+
+例如,麦格拉坦就曾使用遗传算法来搜索极大值的邻域,然后再切换到牛顿法,这一用法在 {doc}`genetic_classifier` 中已有提及。
+
+因此,计量经济学家终究还是采纳了这些适应性算法,只不过并非将其作为他们所研究的主体的*模型*,而是作为自己手中的*工具*。
+
+这份恭维终究得到了回赠,只不过是从侧门进入的:算法跨越了界限,而模型没有。
+
+## 对照追问来解读这份账本
+
+让我们把这两栏账目与我们最初设定的目标对照一番。
+
+**借方**一栏保存着奖赏本身。
+
+一套关于实时转型动态的理论——即东欧改革者们所缺乏的那张地图——到 1993 年为止,在很大程度上仍未被认领,其阻碍在于任意性、需要大量提示、学习任务过于简单,以及经验支持的匮乏。
+
+**贷方**一栏保存着这段旅程沿途所交付的成果:一种在多重均衡中进行选择的、有原则的方法,一种计算那些难以用解析方法处理的均衡的实用方法,以及一次计算技术向计量经济学本身的转移。
+
+贯穿全书的组织性意象,正是计量经济学家与模型内部主体之间的鸿沟。
+
+理性预期通过将主体提升到计量经济学家所不具备的知识水平,来弥合这一鸿沟。
+
+有限理性则提议从另一端来弥合它,把主体降低到计量经济学家的水平,让他们去学习。
+
+按萨金特自己的记账方式,到 1993 年,这一研究计划尚未交出促使它诞生的转型动态理论,而它试图去模仿的计量经济学家们,也一直对这些模型保持距离——理由很充分:这些模型增添了数据无法识别的参数,而计量经济学家想要的恰恰是更少的参数。
+
+但它确实加深了问题的深度,选择了均衡,计算了均衡,并把自己的工具借给了那些婉拒了它的模型的计量经济学家。
+
+这就是 1993 年时,宏观经济学中有限理性前景所处的状态。
+
+## 后记:这一研究计划后来的走向
+
+三十年的时间,足以让我们看清哪些 1993 年的担忧是永久性的,哪些则只是关于一个当时尚未站稳脚跟的领域。
+
+### 任意性被规训了
+
+第一项保留意见是,一旦我们不再坚持主体知晓均衡,就没有任何东西能告诉我们该用什么来取代它。
+
+后来出现的规训机制是**预期稳定性**。
+
+埃文斯和洪卡波亚 {cite:p}`EvansHonkapohja2001` 证明,一个理性预期均衡是否可学习,取决于从感知运动法则到实际运动法则的映射所满足的某个条件——正是 {doc}`bounded_rationality` 中的那个 $T$ 映射——而且这一条件在很大程度上*独立于*学习算法的具体细节。
+
+这正是 1993 年所缺失的东西。
+
+选择不再是建模者恰好写下的那个递归式的偶然产物;一大类合理的学习规则会选出相同的均衡,人们无需进行任何模拟便可判断出这些均衡是什么。
+
+{doc}`olg_adaptive_money` 中的稳定性反转正是一个例证。
+
+在 1993 年,它看起来像是关于最小二乘法的一个特有事实,仅由 {cite:t}`BrunoFischer1990` 用另一种估计方法得出相同结果这一事实来加以支撑。
+
+E-稳定性解释了为何两者会一致。
+
+### 转型动态理论确实到来了,但形式比人们所期望的要狭窄
+
+那份奖赏原本是一套关于非均衡调整的理论。
+
+而这一研究计划所交付的,实际上是一套关于*偏离*均衡的理论:逃逸动态。
+
+我们在 {doc}`olg_adaptive_money` 末尾所模拟出的锯齿形状,在 1993 年不过是一个数值上的趣闻。
+
+{cite:t}`ChoWilliamsSargent2002` 运用大偏差理论对其进行了解析刻画,计算出了最可能的逃逸路径以及逃逸发生的速率,而 {cite:t}`Williams2019` 又大大扩展了这一刻画。
+
+QuantEcon 讲座 {doc}`phillips_escaping_nash` 和 {doc}`phillips_priors` 详细讲解了这两方面的内容。
+
+这比最初追问所要求的要少。
+
+它描述的是从自我确认均衡出发的反复性偏离,而非一个从未有过市场经济的国家迎来市场经济这一事件本身。
+
+但这确实是一套关于一个不会安定下来的系统的、真正的理论,是推导得出而非模拟得出的,而在 1993 年,这样的理论并不存在。
+
+### 计量经济学家终究回礼了
+
+在这一点上,1993 年的评估被彻底超越了。
+
+我们前面看到的障碍,是新增参数存在于一个逐渐消失的瞬态过程之中。
+
+但那一论证仅适用于采用 $1/t$ 增益的学习方案,这种方案会收敛并随后停止移动。
+
+*常数增益学习则不存在这种瞬态过程。* 信念永不安定;它们会持续不断地移动,而它们的运动本身正是数据平稳分布的一部分。
+
+因此,增益像一个普通结构参数那样是可识别的——从整个样本中、以通常的速率识别出来——而不像 $\beta_0$ 那样,只能从一段有界的初始时期中识别。
+
+让我们在我们一直使用的模型上验证这一点。
+
+```{code-cell} ipython3
+def simulate_constant_gain(β0, gain, u):
+ "Bray's cobweb when agents discount old prices at a fixed rate."
+ T = len(u)
+ β, p = np.empty(T), np.empty(T)
+ β[0] = β0
+ for t in range(T):
+ p[t] = a + b * β[t] + u[t]
+ if t + 1 < T:
+ β[t + 1] = β[t] + gain * (p[t] - β[t])
+ return p, β
+
+def ssr_gain(g_hat, p, β0):
+ "Fit criterion for a candidate gain, given observed prices."
+ T = len(p)
+ β = np.empty(T)
+ β[0] = β0
+ for t in range(1, T):
+ β[t] = β[t - 1] + g_hat * (p[t - 1] - β[t - 1])
+ return np.sum((p - a - b * β) ** 2)
+
+gain_true = 0.05
+for T in (500, 5_000, 50_000):
+ penalties = []
+ for seed in range(20): # average out sampling noise
+ u = np.random.default_rng(seed).standard_normal(T)
+ p, _ = simulate_constant_gain(β_star, gain_true, u)
+ penalties.append(ssr_gain(0.08, p, β_star) - ssr_gain(gain_true, p, β_star))
+ print(f"T = {T:6d}: mean penalty for using gain 0.08 instead of 0.05 : "
+ f"{np.mean(penalties):9.1f}")
+```
+
+惩罚值随样本量成比例增长,这与斜率 $b$ 的情形完全一致,而与初始信念的情形恰恰*相反*。
+
+这正是为何最终问世的计量经济学研究使用了常数增益。
+
+{cite:t}`SargentWilliamsZha2006` 利用二战后美国数据,对美联储的常数增益学习模型进行了估计——将一套真正的递归估计程序赋予了模型内部的政府,并询问数据它所使用的是哪一种增益。
+
+{cite:t}`SargentWilliams2005` 研究了政府对漂移系数的先验信念如何塑造其收敛结果,而 {doc}`phillips_priors` 对此有进一步的展开。
+
+{doc}`phillips_drifts_volatilities` 将一个漂移系数向量自回归模型拟合到同一段历史事件上,并追问那究竟是糟糕的政策还是糟糕的运气。
+
+因此,这份恭维终究是双向而行的。
+
+它需要一次学习技术上的转变——从一种会收敛的方案转变为一种永不收敛的方案——才使得这些模型变得可估计,而这一转变是出于经济学而非计量经济学上的考量而做出的。
+
+### 第二次以不同方式进行的撤退
+
+本书中的研究计划保留了个体理性,放弃了相互一致性:主体会进行优化,但所依据的信念仍处在他们自己的估计之中。
+
+一支并行的文献沿着另一条轴线撤退。
+
+在汉森和萨金特 {cite:p}`HansenSargent2008` 的**稳健性**研究中,主体根本不去估计他们的模型。
+
+他们承认自己无法知晓这个模型,于是针对他们无法排除的那些模型中的最坏情形进行优化。
+
+这两种做法是互补的,而非相互竞争的,二者都是对 {doc}`bounded_rationality` 开篇所提问题的回答:我们该如何处理理性预期所赋予的那份知识?
+
+一种回答是,主体应当学习计量经济学家正在学习的东西。
+
+另一种回答是,他们应当在从不学习这些知识的情况下也表现良好。
+
+### 算法持续跨界流动
+
+1993 年的账本记录了,计量经济学家们即便婉拒了这些模型,却仍采纳了这些适应性算法作为计算工具。
+
+这种交流增多了,并再次反转了方向。
+
+{doc}`genetic_classifier` 中霍兰德的"水桶传递"机制——把你的部分回报向后支付给任何促成了你的因素——正是*时序差分学习*,它后来成为现代强化学习的组织性思想 {cite:p}`Sutton_2018`。
+
+事后看来,{doc}`marimon_mcgrattan_sargent` 中的分类器系统,可以被辨认为带有手工构建的函数近似器的强化学习者;而 {doc}`genetic_classifier` 中的感知机,则演变成了 {doc}`back_prop` 中的深度网络,如今正充当着函数近似器的角色。
+
+事实证明,萨金特笔下那些人工智能主体,并非借自邻近领域的一个比喻。
+
+它们是该领域后来所构建成果的一个早期实例。
+
+### 再度解读这份账本
+
+1993 年的借方并未被全部偿清。
+
+选择仍然众多,我们赋予主体的学习任务,相较于真实企业所求解的那些任务,依然显得简单,也没有人拿出东欧改革所呼唤的那套转型动态理论。
+
+但这些条目已经发生了变化。
+
+均衡选择获得了一套理论,而不再仅仅是一组例子;非收敛现象获得了解析刻画,而不再仅仅是模拟结果;而计量经济学家们,一旦获得了一个其额外参数确实能被数据说明的模型版本,便欣然采纳了它。
+
+计量经济学家与模型内部主体之间的鸿沟依然存在。
+
+它变窄了,而我们如今对它的宽度也已经了解了不少。
\ No newline at end of file