2. 调试 Debugging#
概要
调试 Firedrake 程序: 常见报错及其成因, 用 pdb/gdb 调试 Python 代码与生成的 C 代码, 用 -start_in_debugger 和 tmux-mpi 做并行调试, 以及编译 .pyx 文件.
Overview
Debugging Firedrake programs: common errors and their causes, debugging Python and generated C code with pdb/gdb, parallel debugging with -start_in_debugger and tmux-mpi, and compiling .pyx files.
2.1. 常见问题 Common problems#
本节按报错信息组织. 遇到问题时可以直接拿报错里的关键词 (例如 DIVERGED_LINEAR_SOLVE) 在本页搜索.
This section is organized by error message. When a problem shows up, one can search this page directly for a keyword taken from the error (for example DIVERGED_LINEAR_SOLVE).
2.1.1. 线性求解不收敛The linear solve does not converge: DIVERGED_LINEAR_SOLVE#
求解失败时 Firedrake 抛出的异常大致长这样:
When the solve fails, the exception raised by Firedrake looks roughly like this:
File "/home/user/firedrake/src/firedrake/firedrake/adjoint/solving.py", line 50, in wrapper
output = solve(*args, **kwargs)
File "/home/user/firedrake/src/firedrake/firedrake/solving.py", line 129, in solve
_solve_varproblem(*args, **kwargs)
File "/home/user/firedrake/src/firedrake/firedrake/solving.py", line 161, in _solve_varproblem
solver.solve()
File "/home/user/firedrake/src/firedrake/firedrake/adjoint/variational_solver.py", line 75, in wrapper
out = solve(self, **kwargs)
File "/home/user/firedrake/src/firedrake/firedrake/variational_solver.py", line 278, in solve
solving_utils.check_snes_convergence(self.snes)
File "/home/user/firedrake/src/firedrake/firedrake/solving_utils.py", line 139, in check_snes_convergence
raise ConvergenceError(r"""Nonlinear solve failed to converge after %d nonlinear iterations.
firedrake.exceptions.ConvergenceError: Nonlinear solve failed to converge after 0 nonlinear iterations.
Reason:
DIVERGED_LINEAR_SOLVE
这条信息只说明”线性求解没有成功”, 并没有说为什么. 常见成因有:
方程没有定好. 最常见的是边界条件写错或写漏, 使离散系统奇异或严重病态. 先逐条核对
DirichletBC的边界编号和取值.系统本身是奇异的. 例如纯 Neumann 问题, 解只差一个常数. 这类问题要给出零空间 (
nullspace), 或者用 Lagrange 乘子把常数固定下来.迭代次数用完了. 迭代法不收敛未必是方程错了, 也可能只是预条件子不合适, 或者
ksp_max_it给得太小.外部库报错. 用直接法 (LU) 时, 错误常常来自 PETSc 调用的 MUMPS 等外部库, 见下面两个例子.
排查步骤. Firedrake 抛出的异常只到”线性求解失败”为止, 真正的原因要到 PETSc 的 KSP / PC 层去问 (取用方法见 PETSc 一章的 KSP 一节).
This message only says that “the linear solve did not succeed”, not why. The common causes are:
The equation is not set up correctly. Most often a boundary condition is wrong or missing, which makes the discrete system singular or severely ill-conditioned. Check the boundary ids and the values of every
DirichletBCone by one first.The system itself is singular. For example a pure Neumann problem, whose solution is determined only up to a constant. Such problems need a nullspace (
nullspace), or a Lagrange multiplier to pin the constant down.The iterations ran out. A non-converging iterative method does not necessarily mean the equation is wrong; the preconditioner may simply be unsuitable, or
ksp_max_itmay be too small.An external library reported an error. With a direct method (LU) the error often comes from an external library called by PETSc such as MUMPS, see the two examples below.
Checking steps. The exception raised by Firedrake stops at “the linear solve failed”; the real cause has to be asked for at PETSc’s KSP / PC level (for how to get hold of them see the section KSP in the PETSc chapter).
让 PETSc 自己报告残差历史和收敛原因:Let PETSc itself report the residual history and the convergence reason:
solver_parameters = { # ... 原有的参数 'ksp_monitor': None, 'ksp_converged_reason': None, }
输出形如 (下面是把
ksp_max_it设成 2 故意制造的不收敛):The output looks like this (below, the non-convergence is produced on purpose by settingksp_max_itto 2):Residual norms for firedrake_1_ solve. 0 KSP Residual norm 6.091695533358e-02 1 KSP Residual norm 2.802120913580e-03 2 KSP Residual norm 1.265467409231e-03 Linear firedrake_1_ solve did not converge due to DIVERGED_ITS iterations 2DIVERGED_ITS才是真正的原因 (迭代次数用完), 而 Firedrake 的异常里只有DIVERGED_LINEAR_SOLVE.DIVERGED_ITSis the real reason (the iterations ran out), while Firedrake’s exception only carriesDIVERGED_LINEAR_SOLVE.加
'ksp_error_if_not_converged': None, 让 PETSc 在失败的地方直接抛错, 异常里会带上 C 端的调用栈:Add'ksp_error_if_not_converged': Noneto make PETSc raise the error where it fails; the exception then carries the call stack of the C side:petsc4py.PETSc.Error: error code 91 [0] SNESSolve() at .../src/snes/interface/snes.c:4875 [0] SNESSolve_KSPONLY() at .../src/snes/impls/ksponly/ksponly.c:49 [0] KSPSolve() at .../src/ksp/ksp/interface/itfunc.c:1092 [0] KSPSolve_Private() at .../src/ksp/ksp/interface/itfunc.c:999 [0] KSPSolve() has not converged, reason DIVERGED_ITS
错误码 91 就是
PETSC_ERR_NOT_CONVERGED, 对照表见后面的PetscErrorCode一节.Error code 91 isPETSC_ERR_NOT_CONVERGED; the table is given in thePetscErrorCodesection below.加
'ksp_view': None确认实际用的求解器和预条件子是不是你以为的那个. 注意 Firedrake 的默认求解器是直接法 (preonly+lu), 不显式设置ksp_type/pc_type时并不会真的做迭代.Add'ksp_view': Noneto confirm that the solver and the preconditioner actually used are the ones you think they are. Note that Firedrake’s default solver is a direct method (preonly+lu), so unlessksp_type/pc_typeare set explicitly no iteration is performed at all.
下面是两个来自 MUMPS 的实际例子.
Below are two real examples coming from MUMPS.
示例 1: MUMPS 工作空间不足 (INFOG(1)=-9)Example 1: MUMPS runs out of workspace (INFOG(1)=-9)
用直接法时, 错误往往来自 PETSc 调用的外部库. 下面这段是 MUMPS 在数值分解阶段报的错, 关键是 INFOG(1) 和 INFO(2) 两个数字, 拿它们去查 MUMPS 手册即可.
参考资料:
MUMPS 用户手册: https://graal.ens-lyon.fr/MUMPS/index.php?page=doc
PETSc 的
MATSOLVERMUMPS手册页: https://petsc.org/main/manualpages/Mat/MATSOLVERMUMPS/
With a direct method the error often comes from an external library called by PETSc. The block below is an error reported by MUMPS in the numerical factorization phase; what matters are the two numbers INFOG(1) and INFO(2), which are then looked up in the MUMPS manual.
References:
The MUMPS user’s guide: https://graal.ens-lyon.fr/MUMPS/index.php?page=doc
PETSc’s
MATSOLVERMUMPSmanual page: https://petsc.org/main/manualpages/Mat/MATSOLVERMUMPS/
[63]PETSC ERROR: --------------------- Error Message --------------------------------------------------------------
[63]PETSC ERROR: Error in external library
[63]PETSC ERROR: Error reported by MUMPS in numerical factorization phase: INFOG(1)=-9, INFO(2)=26
[63]PETSC ERROR: See https://petsc.org/release/faq/ for trouble shooting.
[63]PETSC ERROR: Petsc Development GIT revision: v3.4.2-38777-g979cc68729 GIT Date: 2022-06-22 21:18:23 +0100
[63]PETSC ERROR: ../hsolver/hsolver.py on a default named myhost by user Tue Nov 8 16:03:18 2022
[63]PETSC ERROR: Configure options PETSC_DIR=/home/user/opt/firedrake-env/firedrake-complex-int64/src/petsc PETSC_ARCH=default --download-ptscotch --with-zlib --download-hwloc --with-c2html=0 --download-eigen="/home/user/opt/firedrake-env/firedrake-complex-int64/src/eigen-3.3.3.tgz " --download-mpich --download-hdf5 --with-fortran-bindings=0 --with-64-bit-indices --download-bison --with-cxx-dialect=C++11 --download-metis --download-openblas --download-openblas-make-options="'USE_THREAD=0 USE_LOCKING=1 USE_OPENMP=0'" --download-pastix --download-mumps --with-shared-libraries=1 --with-scalar-type=complex --download-cmake --download-scalapack --with-debugging=0 --download-netcdf --download-superlu_dist --download-suitesparse --download-pnetcdf
[63]PETSC ERROR: #1 MatFactorNumeric_MUMPS() at /home/user/opt/firedrake-env/firedrake-complex-int64/src/petsc/src/mat/impls/aij/mpi/mumps/mumps.c:1664
[63]PETSC ERROR: #2 MatLUFactorNumeric() at /home/user/opt/firedrake-env/firedrake-complex-int64/src/petsc/src/mat/interface/matrix.c:3177
[63]PETSC ERROR: #3 PCSetUp_LU() at /home/user/opt/firedrake-env/firedrake-complex-int64/src/petsc/src/ksp/pc/impls/factor/lu/lu.c:135
[63]PETSC ERROR: #4 PCSetUp() at /home/user/opt/firedrake-env/firedrake-complex-int64/src/petsc/src/ksp/pc/interface/precon.c:993
[63]PETSC ERROR: #5 KSPSetUp() at /home/user/opt/firedrake-env/firedrake-complex-int64/src/petsc/src/ksp/ksp/interface/itfunc.c:407
[63]PETSC ERROR: #6 KSPSolve_Private() at /home/user/opt/firedrake-env/firedrake-complex-int64/src/petsc/src/ksp/ksp/interface/itfunc.c:843
[63]PETSC ERROR: #7 KSPSolve() at /home/user/opt/firedrake-env/firedrake-complex-int64/src/petsc/src/ksp/ksp/interface/itfunc.c:1078
MUMPS 手册中关于 INFOG(1) = -9 和 ICNTL(14) 的说明:
What the MUMPS manual says about INFOG(1) = -9 and ICNTL(14):
The main internal real/complex workarray S is too small. If INFO(2) is positive, then the number
of entries that are missing in S at the moment when the error is raised is available in INFO(2).
If INFO(2) is negative, then its absolute value should be multiplied by 1 million. If an error –9
occurs, the user should increase the value of ICNTL(14) before calling the factorization (JOB=2)
again, except if LWK_USER is provided LWK_USER should be increased.
ICNTL(14) corresponds to the percentage increase in the estimated working space.
Phase: accessed by the host both during the analysis and the factorization phases.
Default value: between 20 and 35 (which corresponds to at most 35 % increase) and depends on
the number of MPI processes. It is set to 5 % with SYM=1 and one MPI process.
Related parameters: ICNTL(23)
Remarks: When significant extra fill-in is caused by numerical pivoting, increasing ICNTL(14)
may help
也就是说 -9 表示 MUMPS 内部的实数工作数组不够大, 办法是把 ICNTL(14) (估计工作空间的放大百分比) 调大之后重新分解.
MUMPS 的参数在 PETSc 里通过 -mat_mumps_icntl_<num> 设置, 写进 Firedrake 的 solver_parameters 时要放在顶层:
In other words, -9 means that the internal real workarray of MUMPS is too small, and the remedy is to increase ICNTL(14) (the percentage increase of the estimated working space) and factorize again.
The MUMPS parameters are set in PETSc through -mat_mumps_icntl_<num>; when written into Firedrake’s solver_parameters they have to be placed at the top level:
solver_parameters = {
'ksp_type': 'preonly',
'pc_type': 'lu',
'pc_factor_mat_solver_type': 'mumps',
'mat_mumps_icntl_14': 200, # 工作空间放大 200%
}
Tip
求解器参数拼错了不会报错, 只会被悄悄忽略. 想确认某个 MUMPS 参数真的生效了, 加上 'mat_mumps_icntl_4': 2 让 MUMPS 把它实际使用的 ICNTL 值打印出来:
ICNTL(14) Percentage of memory relaxation = 20
本机实测 (Firedrake 2026.4.1, MUMPS 5.8.2): 顶层的 mat_mumps_icntl_14 能改动这一行的数值, 而写成 pc_factor_mat_mumps_icntl_14 时打印出来的仍是 MUMPS 自己的默认值 20.
Tip
A misspelled solver parameter raises no error and is silently ignored. To confirm that a MUMPS parameter really took effect, add 'mat_mumps_icntl_4': 2 so that MUMPS prints the ICNTL values it actually uses:
ICNTL(14) Percentage of memory relaxation = 20
Measured on this machine (Firedrake 2026.4.1, MUMPS 5.8.2): mat_mumps_icntl_14 at the top level does change the number on this line, whereas written as pc_factor_mat_mumps_icntl_14 the printed value is still MUMPS’ own default of 20.
其他有用的选项Other useful options
ksp_view、ksp_monitor、ksp_converged_reason、ksp_error_if_not_converged.ksp_view, ksp_monitor, ksp_converged_reason, ksp_error_if_not_converged.
示例 2: MUMPS 遇到零主元 (INFOG(1)=-10)Example 2: MUMPS hits a zero pivot (INFOG(1)=-10)
下面这条是 petsc4py 直接抛出来的 PETSc 错误, 错误码 76 即 PETSC_ERR_LIB, 表示 PETSc 调用的外部库出了错:
The following is a PETSc error raised directly by petsc4py; error code 76 is PETSC_ERR_LIB, meaning that an external library called by PETSc failed:
petsc4py.PETSc.Error: error code 76
[0] SNESSolve() at /home/user/firedrake/real-int32-mkl-debug/src/petsc/src/snes/interface/snes.c:4693
[0] SNESSolve_KSPONLY() at /home/user/firedrake/real-int32-mkl-debug/src/petsc/src/snes/impls/ksponly/ksponly.c:48
[0] KSPSolve() at /home/user/firedrake/real-int32-mkl-debug/src/petsc/src/ksp/ksp/interface/itfunc.c:1070
[0] KSPSolve_Private() at /home/user/firedrake/real-int32-mkl-debug/src/petsc/src/ksp/ksp/interface/itfunc.c:824
[0] KSPSetUp() at /home/user/firedrake/real-int32-mkl-debug/src/petsc/src/ksp/ksp/interface/itfunc.c:405
[0] PCSetUp() at /home/user/firedrake/real-int32-mkl-debug/src/petsc/src/ksp/pc/interface/precon.c:994
[0] PCSetUp_LU() at /home/user/firedrake/real-int32-mkl-debug/src/petsc/src/ksp/pc/impls/factor/lu/lu.c:120
[0] MatLUFactorNumeric() at /home/user/firedrake/real-int32-mkl-debug/src/petsc/src/mat/interface/matrix.c:3215
[0] MatFactorNumeric_MUMPS() at /home/user/firedrake/real-int32-mkl-debug/src/petsc/src/mat/impls/aij/mpi/mumps/mumps.c:1683
[0] Error in external library
[0] Error reported by MUMPS in numerical factorization phase: INFOG(1)=-10, INFO(2)=4464
MUMPS 手册对 INFOG(1)=-10 的说明:
What the MUMPS manual says about INFOG(1)=-10:
INFOG(1)=-10: Numerically singular matrix. INFO(2) holds the number of eliminated pivots.
也就是矩阵在数值上是奇异的. 打开 MUMPS 的零主元检测可以让分解不再直接失败:
So the matrix is numerically singular. Turning on MUMPS’ zero pivot detection keeps the factorization from failing outright:
solver_parameters = {
'ksp_type': 'preonly',
'pc_type': 'lu',
'pc_factor_mat_solver_type': 'mumps',
'mat_mumps_icntl_24': 1, # 检测零主元
# 'mat_mumps_cntl_3': 1e-12, # 判定零主元的阈值
}
ICNTL(24) controls the detection of ``null pivot rows''.
Phase: accessed by the host during the factorization phase
Possible variables/arrays involved: PIVNUL LIST
Possible values :
0: Nothing done. A null pivot row will result in error INFO(1)=-10.
1: Null pivot row detection.
Other values are treated as 0.
Default value: 0 (no null pivot row detection)
Warning
ICNTL(24) 只是让分解阶段不因为零主元而报错, 它并不替你处理奇异性: 解在零空间方向上的分量是任意的.
如果矩阵奇异是问题本身决定的 (例如纯 Neumann 问题), 正确的做法是给 solve 传 nullspace. 用直接法时两件事都要做: nullspace 只负责把零空间分量投影掉, 不参与 LU 分解, 而 ICNTL(24) 才让分解本身能走完.
Warning
ICNTL(24) only keeps the factorization phase from erroring out on a zero pivot; it does not deal with the singularity for you: the component of the solution along the nullspace is arbitrary.
If the singularity is inherent to the problem (a pure Neumann problem, say), the right thing to do is to pass a nullspace to solve. With a direct method both are needed: nullspace only projects the nullspace component out and takes no part in the LU factorization, while ICNTL(24) is what lets the factorization itself run to completion.
2.1.2. 高阶网格上对法向求导: ReferenceGrad 报错 Differentiating along the normal on a higher-order mesh: the ReferenceGrad error#
在比较老的 Firedrake/UFL 上, 对高阶网格 (等参元) 的单元边界外法向求导会报下面的错:
On older versions of Firedrake/UFL, differentiating with respect to the outward normal on the facets of a higher-order (isoparametric) mesh raised the following error:
ufl.log.UFLException: Currently no support for ReferenceGrad in CoefficientDerivative.
这条限制来自 UFL: 对系数求导时, 它不支持表达式里出现参考单元上的梯度 ReferenceGrad. 高阶网格上的 FacetNormal 依赖坐标场, 求导时就可能落到这条分支上.
下面的单元把两种写法各在仿射网格和高阶网格上跑一遍, 用来检查当前版本还有没有这个限制: get_s12 先对 u 求三次导再和法向缩并, get_s12_v2 则多了一次”对含法向的表达式再求导”.
The restriction comes from UFL: when differentiating with respect to a coefficient it does not support the gradient on the reference cell, ReferenceGrad, appearing in the expression. On a higher-order mesh FacetNormal depends on the coordinate field, so differentiating it may end up on this branch.
The cell below runs both formulations on an affine mesh and on a higher-order mesh, so as to check whether the current version still has this restriction: get_s12 takes three derivatives of u and then contracts them with the normal, while get_s12_v2 additionally differentiates an expression that already contains the normal.
Note
本机实测 (Firedrake 2026.4.1, UFL 2025.3.0): 这四种组合现在全部通过, 单元输出四行 TEST OK!, 这条报错在这条路径上已经不再出现.
UFL 里这条限制本身还在, 只是异常类型变了: 抛出点在 ufl/algorithms/apply_derivatives.py 的 ReferenceGrad 分支, 现在抛的是 NotImplementedError 而不是 UFLException (ufl.log 模块已经没有了). 真遇到时, 按 Currently no support for ReferenceGrad in CoefficientDerivative 这句原文去搜索仍然有效.
下面的单元保留下来当回归检查用: 哪天它打印出 TEST ERROR, 就说明这条限制又回来了.
Note
Measured on this machine (Firedrake 2026.4.1, UFL 2025.3.0): all four combinations now pass, the cell prints four lines of TEST OK!, and this error no longer shows up on this path.
The restriction itself is still present in UFL, only the exception type has changed: it is raised in the ReferenceGrad branch of ufl/algorithms/apply_derivatives.py, and what is raised now is a NotImplementedError rather than a UFLException (the ufl.log module is gone). Should one run into it, searching for the original wording Currently no support for ReferenceGrad in CoefficientDerivative still works.
The cell below is kept as a regression check: the day it prints TEST ERROR, the restriction is back.
from firedrake import *
def get_s12(mesh, u, v):
n = FacetNormal(mesh)
s1 = dot(n, dot(n, grad(grad(grad(u)))))
s2 = dot(n, dot(n, grad(grad(grad(v)))))
return s1, s2
def get_s12_v2(mesh, u, v):
n = FacetNormal(mesh)
s1 = dot(n, grad(dot(n, grad(grad(u)))))
s2 = dot(n, grad(dot(n, grad(grad(v)))))
return s1, s2
def test_grad_n(high_order_mesh=True, fun=get_s12):
p = 2
N = 4
mesh = RectangleMesh(N, N, 1, 1)
if high_order_mesh:
V = VectorFunctionSpace(mesh, 'CG', 2)
coords = Function(V).interpolate(mesh.coordinates)
mesh = Mesh(coords)
V = FunctionSpace(mesh, 'CG', p)
u, v = TrialFunction(V), TestFunction(V)
s1, s2 = fun(mesh, u, v)
a = inner(grad(u), grad(v))*dx + inner(s1('+'), s2('+'))*dS + inner(s1('-'), s2('-'))*dS
L = inner(Constant(0), v)*dx
sol = Function(V)
prob = LinearVariationalProblem(a, L, sol)
solver = LinearVariationalSolver(prob)
solver.solve()
for f in [get_s12, get_s12_v2]:
for hom in [False, True]:
try:
test_grad_n(high_order_mesh=hom, fun=f)
print(f.__name__, f': High order mesh: {hom}, ', "TEST OK!")
except Exception as e:
print(f.__name__, f': High order mesh: {hom}, ', "TEST ERROR: ", e)
get_s12 : High order mesh: False, TEST OK!
get_s12 : High order mesh: True, TEST OK!
get_s12_v2 : High order mesh: False, TEST OK!
get_s12_v2 : High order mesh: True, TEST OK!
2.1.3. PyErr_Occurred#
python: src/petsc4py.PETSc.c:348918: __Pyx_PyCFunction_FastCall: Assertion `!PyErr_Occurred()' failed.
这类断言失败通常发生在 PETSc 回调 Python 代码的时候: 你写的回调 (自定义预条件子、KSP/SNES 的 monitor、PCPython 等) 里抛了异常, 异常没有被正确地传回 C 端, 于是下一次进入 petsc4py 时断言 !PyErr_Occurred() 就失败了. 报错位置在 petsc4py 生成的 C 文件里, 和你的代码没有直接对应关系, 看那个行号没有用.
排查办法: 把回调函数单独拿出来在 Python 里调用一次, 先保证它自己不出错 —— 最常见的就是变量名拼错、变量没定义这类低级错误. 再配合下一节让 PETSc 打印完整错误信息, 通常能定位到更靠前的出错点.
Assertion failures of this kind usually happen when PETSc calls back into Python code: the callback you wrote (a custom preconditioner, a monitor of KSP/SNES, PCPython and so on) raised an exception, the exception was not propagated back to the C side properly, and so the assertion !PyErr_Occurred() fails the next time petsc4py is entered. The reported location is inside the C file generated by petsc4py and has no direct correspondence to your code, so the line number there is of no use.
How to track it down: take the callback out and call it once from Python on its own, making sure it does not fail by itself — the most common causes are trivial ones such as a misspelled or undefined variable name. Combined with the next section, which makes PETSc print the full error message, this usually pins down an earlier point of failure.
2.1.4. 让 PETSc 打印完整的错误信息 Making PETSc print the full error message#
petsc4py 默认会把 PETSc 的错误压缩成一个 Python 异常, 只保留几行调用栈. 在程序最开始加上
By default petsc4py compresses a PETSc error into a Python exception and keeps only a few lines of the call stack. Adding
from firedrake.petsc import PETSc
PETSc.Sys.popErrorHandler()
就会退回 PETSc 自带的错误处理器, 出错时把完整的 [0]PETSC ERROR: 块打印到标准错误, 里面有 PETSc 版本、configure 选项和 C 端逐层调用栈. 向 Firedrake 或 PETSc 提交 issue 时, 对方需要的正是这一整块信息.
本机实测同一段出错代码, 加与不加的区别:
at the very beginning of the program falls back to PETSc’s own error handler, which prints the complete [0]PETSC ERROR: block to standard error when something goes wrong; it contains the PETSc version, the configure options and the call stack of the C side level by level. When filing an issue against Firedrake or PETSc, this whole block is exactly what is needed.
Measured on this machine with the same failing code, the difference between leaving it out and putting it in (the first block is without it — the petsc4py default; the second is with PETSc.Sys.popErrorHandler()):
# 不加 (petsc4py 默认): 只有压缩过的几行
petsc4py.PETSc.Error: error code 56
[0] VecView() at .../src/vec/vec/interface/vector.c:894
[0] No support for this operation for this object type
[0] No method view for Vec of type (null)
# 加了 PETSc.Sys.popErrorHandler(): PETSc 自己的完整报告
[0]PETSC ERROR: --------------------- Error Message ------------------------------
[0]PETSC ERROR: No support for this operation for this object type
[0]PETSC ERROR: No method view for Vec of type (null)
[0]PETSC ERROR: See https://petsc.org/release/faq/ for trouble shooting.
[0]PETSC ERROR: PETSc Release Version 3.25.0, unknown
[0]PETSC ERROR: <脚本名> with 1 MPI process(es) and PETSC_ARCH ... on <主机名> by <用户名> <时间>
[0]PETSC ERROR: Configure options: ...
[0]PETSC ERROR: #1 VecView() at .../src/vec/vec/interface/vector.c:894
2.1.5. PetscErrorCode#
petsc4py 抛出的 PETSc.Error 带一个 ierr 属性, 就是 PETSc 的错误码; 前面出现的 error code 91、error code 76、error code 56 都是它. 拿到错误码之后对照下表就知道属于哪一类, 常用的几个:
错误码 |
名字 |
含义 |
|---|---|---|
56 |
|
该对象类型不支持这个操作 (例如对还没设类型的 |
63 |
|
参数越界 |
71 |
|
LU 分解遇到零主元 |
76 |
|
PETSc 调用的外部库 (MUMPS、hypre 等) 出错 |
91 |
|
求解器没有收敛 |
在代码里可以按 ierr 分情况处理, 完整例子见 PETSc 一章的 KSP 一节.
手册页: https://petsc.org/release/manualpages/Sys/PetscErrorCode/
The PETSc.Error raised by petsc4py carries an ierr attribute, which is PETSc’s error code; the error code 91, error code 76 and error code 56 seen above are all of this kind. Once the code is known, the table below says which class it belongs to; the common ones are:
Error code |
Name |
Meaning |
|---|---|---|
56 |
|
the operation is not supported for this object type (for example calling |
63 |
|
an argument is out of range |
71 |
|
a zero pivot was hit during the LU factorization |
76 |
|
an external library called by PETSc (MUMPS, hypre and so on) failed |
91 |
|
the solver did not converge |
The code can branch on ierr; for a complete example see the section KSP in the PETSc chapter.
Manual page: https://petsc.org/release/manualpages/Sys/PetscErrorCode/
完整清单 (PETSc 3.25 的 include/petscsystypes.h)
{
PETSC_SUCCESS = 0,
PETSC_ERR_BOOLEAN_MACRO_FAILURE = 1, /* do not use */
PETSC_ERR_MIN_VALUE = 54, /* should always be one less than the smallest value */
PETSC_ERR_MEM = 55, /* unable to allocate requested memory */
PETSC_ERR_SUP = 56, /* no support for requested operation */
PETSC_ERR_SUP_SYS = 57, /* no support for requested operation on this computer system */
PETSC_ERR_ORDER = 58, /* operation done in wrong order */
PETSC_ERR_SIG = 59, /* signal received */
PETSC_ERR_FP = 72, /* floating point exception */
PETSC_ERR_COR = 74, /* corrupted PETSc object */
PETSC_ERR_LIB = 76, /* error in library called by PETSc */
PETSC_ERR_PLIB = 77, /* PETSc library generated inconsistent data */
PETSC_ERR_MEMC = 78, /* memory corruption */
PETSC_ERR_CONV_FAILED = 82, /* iterative method (KSP or SNES) failed */
PETSC_ERR_USER = 83, /* user has not provided needed function */
PETSC_ERR_SYS = 88, /* error in system call */
PETSC_ERR_POINTER = 70, /* pointer does not point to valid address */
PETSC_ERR_MPI_LIB_INCOMP = 87, /* MPI library at runtime is not compatible with MPI user compiled with */
PETSC_ERR_ARG_SIZ = 60, /* nonconforming object sizes used in operation */
PETSC_ERR_ARG_IDN = 61, /* two arguments not allowed to be the same */
PETSC_ERR_ARG_WRONG = 62, /* wrong argument (but object probably ok) */
PETSC_ERR_ARG_CORRUPT = 64, /* null or corrupted PETSc object as argument */
PETSC_ERR_ARG_OUTOFRANGE = 63, /* input argument, out of range */
PETSC_ERR_ARG_BADPTR = 68, /* invalid pointer argument */
PETSC_ERR_ARG_NOTSAMETYPE = 69, /* two args must be same object type */
PETSC_ERR_ARG_NOTSAMECOMM = 80, /* two args must be same communicators */
PETSC_ERR_ARG_WRONGSTATE = 73, /* object in argument is in wrong state, e.g. unassembled mat */
PETSC_ERR_ARG_TYPENOTSET = 89, /* the type of the object has not yet been set */
PETSC_ERR_ARG_INCOMP = 75, /* two arguments are incompatible */
PETSC_ERR_ARG_NULL = 85, /* argument is null that should not be */
PETSC_ERR_ARG_UNKNOWN_TYPE = 86, /* type name doesn't match any registered type */
PETSC_ERR_FILE_OPEN = 65, /* unable to open file */
PETSC_ERR_FILE_READ = 66, /* unable to read from file */
PETSC_ERR_FILE_WRITE = 67, /* unable to write to file */
PETSC_ERR_FILE_UNEXPECTED = 79, /* unexpected data in file */
PETSC_ERR_MAT_LU_ZRPVT = 71, /* detected a zero pivot during LU factorization */
PETSC_ERR_MAT_CH_ZRPVT = 81, /* detected a zero pivot during Cholesky factorization */
PETSC_ERR_INT_OVERFLOW = 84,
PETSC_ERR_FLOP_COUNT = 90,
PETSC_ERR_NOT_CONVERGED = 91, /* solver did not converge */
PETSC_ERR_MISSING_FACTOR = 92, /* MatGetFactor() failed */
PETSC_ERR_OPT_OVERWRITE = 93, /* attempted to over write options which should not be changed */
PETSC_ERR_WRONG_MPI_SIZE = 94, /* example/application run with number of MPI ranks it does not support */
PETSC_ERR_USER_INPUT = 95, /* missing or incorrect user input */
PETSC_ERR_GPU_RESOURCE = 96, /* unable to load a GPU resource, for example cuBLAS */
PETSC_ERR_GPU = 97, /* An error from a GPU call, this may be due to lack of resources on the GPU or a true error in the call */
PETSC_ERR_MPI = 98, /* general MPI error */
PETSC_ERR_RETURN = 99, /* PetscError() incorrectly returned an error code of 0 */
PETSC_ERR_MEM_LEAK = 100, /* memory alloc/free imbalance */
PETSC_ERR_PYTHON = 101, /* Exception in Python */
PETSC_ERR_MAX_VALUE = 102, /* this is always the one more than the largest error code */
/*
do not use, exist purely to make the enum bounds equal that of a regular int (so conversion
to int in main() is not undefined behavior)
*/
PETSC_ERR_MIN_SIGNED_BOUND_DO_NOT_USE = INT_MIN,
PETSC_ERR_MAX_SIGNED_BOUND_DO_NOT_USE = INT_MAX
} PETSC_ERROR_CODE_ENUM_NAME;
Full list (from include/petscsystypes.h of PETSc 3.25)
{
PETSC_SUCCESS = 0,
PETSC_ERR_BOOLEAN_MACRO_FAILURE = 1, /* do not use */
PETSC_ERR_MIN_VALUE = 54, /* should always be one less than the smallest value */
PETSC_ERR_MEM = 55, /* unable to allocate requested memory */
PETSC_ERR_SUP = 56, /* no support for requested operation */
PETSC_ERR_SUP_SYS = 57, /* no support for requested operation on this computer system */
PETSC_ERR_ORDER = 58, /* operation done in wrong order */
PETSC_ERR_SIG = 59, /* signal received */
PETSC_ERR_FP = 72, /* floating point exception */
PETSC_ERR_COR = 74, /* corrupted PETSc object */
PETSC_ERR_LIB = 76, /* error in library called by PETSc */
PETSC_ERR_PLIB = 77, /* PETSc library generated inconsistent data */
PETSC_ERR_MEMC = 78, /* memory corruption */
PETSC_ERR_CONV_FAILED = 82, /* iterative method (KSP or SNES) failed */
PETSC_ERR_USER = 83, /* user has not provided needed function */
PETSC_ERR_SYS = 88, /* error in system call */
PETSC_ERR_POINTER = 70, /* pointer does not point to valid address */
PETSC_ERR_MPI_LIB_INCOMP = 87, /* MPI library at runtime is not compatible with MPI user compiled with */
PETSC_ERR_ARG_SIZ = 60, /* nonconforming object sizes used in operation */
PETSC_ERR_ARG_IDN = 61, /* two arguments not allowed to be the same */
PETSC_ERR_ARG_WRONG = 62, /* wrong argument (but object probably ok) */
PETSC_ERR_ARG_CORRUPT = 64, /* null or corrupted PETSc object as argument */
PETSC_ERR_ARG_OUTOFRANGE = 63, /* input argument, out of range */
PETSC_ERR_ARG_BADPTR = 68, /* invalid pointer argument */
PETSC_ERR_ARG_NOTSAMETYPE = 69, /* two args must be same object type */
PETSC_ERR_ARG_NOTSAMECOMM = 80, /* two args must be same communicators */
PETSC_ERR_ARG_WRONGSTATE = 73, /* object in argument is in wrong state, e.g. unassembled mat */
PETSC_ERR_ARG_TYPENOTSET = 89, /* the type of the object has not yet been set */
PETSC_ERR_ARG_INCOMP = 75, /* two arguments are incompatible */
PETSC_ERR_ARG_NULL = 85, /* argument is null that should not be */
PETSC_ERR_ARG_UNKNOWN_TYPE = 86, /* type name doesn't match any registered type */
PETSC_ERR_FILE_OPEN = 65, /* unable to open file */
PETSC_ERR_FILE_READ = 66, /* unable to read from file */
PETSC_ERR_FILE_WRITE = 67, /* unable to write to file */
PETSC_ERR_FILE_UNEXPECTED = 79, /* unexpected data in file */
PETSC_ERR_MAT_LU_ZRPVT = 71, /* detected a zero pivot during LU factorization */
PETSC_ERR_MAT_CH_ZRPVT = 81, /* detected a zero pivot during Cholesky factorization */
PETSC_ERR_INT_OVERFLOW = 84,
PETSC_ERR_FLOP_COUNT = 90,
PETSC_ERR_NOT_CONVERGED = 91, /* solver did not converge */
PETSC_ERR_MISSING_FACTOR = 92, /* MatGetFactor() failed */
PETSC_ERR_OPT_OVERWRITE = 93, /* attempted to over write options which should not be changed */
PETSC_ERR_WRONG_MPI_SIZE = 94, /* example/application run with number of MPI ranks it does not support */
PETSC_ERR_USER_INPUT = 95, /* missing or incorrect user input */
PETSC_ERR_GPU_RESOURCE = 96, /* unable to load a GPU resource, for example cuBLAS */
PETSC_ERR_GPU = 97, /* An error from a GPU call, this may be due to lack of resources on the GPU or a true error in the call */
PETSC_ERR_MPI = 98, /* general MPI error */
PETSC_ERR_RETURN = 99, /* PetscError() incorrectly returned an error code of 0 */
PETSC_ERR_MEM_LEAK = 100, /* memory alloc/free imbalance */
PETSC_ERR_PYTHON = 101, /* Exception in Python */
PETSC_ERR_MAX_VALUE = 102, /* this is always the one more than the largest error code */
/*
do not use, exist purely to make the enum bounds equal that of a regular int (so conversion
to int in main() is not undefined behavior)
*/
PETSC_ERR_MIN_SIGNED_BOUND_DO_NOT_USE = INT_MIN,
PETSC_ERR_MAX_SIGNED_BOUND_DO_NOT_USE = INT_MAX
} PETSC_ERROR_CODE_ENUM_NAME;
2.2. 调试 Python 代码 Debugging Python code#
程序抛异常时, 先把回溯的最后几行读完: 找到出错的那一行, 再检查相关变量是不是有异常值 (None、形状不对、nan).
在 Jupyter notebook 里, 出错之后直接执行
When the program raises an exception, first read the last few lines of the traceback to the end: find the line that failed, then check whether the variables involved have unexpected values (None, a wrong shape, nan).
In a Jupyter notebook, running
%debug
就会在出错的那一帧打开 pdb 调试器: p 变量名 打印变量, u / d 在调用栈上下移动, q 退出. 如果希望以后一出错就自动进调试器, 执行一次 %pdb on 即可.
写成脚本运行时对应的做法是: 在可疑位置插一行 breakpoint(), 或者用
right after the error opens the pdb debugger in the frame that failed: p <variable> prints a variable, u / d move up and down the call stack, q quits. To enter the debugger automatically on every future error, run %pdb on once.
The counterpart when running as a script is to insert a breakpoint() line at the suspicious place, or to use
python -m pdb -c continue test.py
让脚本一直跑到出错处再停下.
Firedrake 里有一类异常要特别注意: 报错是 petsc4py.PETSc.Error, Python 回溯只到调用 PETSc 的那一行为止, 真正的原因在 PETSc 的 C 端. 这时 Python 调试器帮不上忙, 应当先按上一节的办法让 PETSc 打印完整错误信息, 不够再上 gdb.
so that the script runs until it fails and stops there.
One class of exceptions in Firedrake deserves particular attention: the error is a petsc4py.PETSc.Error, the Python traceback stops at the line that called PETSc, and the real cause is on PETSc’s C side. A Python debugger is of no help here; one should first make PETSc print the full error message as described in the previous section, and only reach for gdb if that is not enough.
2.3. 调试 C 代码 Debugging C code (gdb)#
Firedrake 的网格管理和线性方程组求解都是 PETSc 做的, 出错也可能出在 PETSc 的 C 代码里. 需要 gdb 这类工具的判据是: 进程死掉或者卡住, 但没有留下 Python 回溯. 典型的有
段错误 (
SIGSEGV)、总线错误 (SIGBUS) 之类的信号, 进程直接被杀掉;程序停在那里不动 (并行时最常见, 见 并行常见陷阱 Common parallel pitfalls), 需要看清它卡在哪个函数上;
C 端的断言失败, 例如前面的
PyErr_Occurred.
反过来, 如果 PETSc 是通过错误码返回的, petsc4py 会把它转成 Python 异常, 这时用不上 gdb. 例如下面这个脚本:
Mesh management and the solution of linear systems in Firedrake are both done by PETSc, so a failure may also occur inside PETSc’s C code. A tool such as gdb is needed when the process dies or hangs without leaving a Python traceback. Typical cases are
signals such as a segmentation fault (
SIGSEGV) or a bus error (SIGBUS), which kill the process outright;the program sitting there doing nothing (most common in parallel, see 并行常见陷阱 Common parallel pitfalls), where one needs to see which function it is stuck in;
an assertion failure on the C side, such as the
PyErr_Occurredabove.
Conversely, when PETSc returns through an error code, petsc4py turns it into a Python exception and gdb is not needed. Take the following script:
# filename: test.py
import sys
import petsc4py
petsc4py.init(sys.argv)
from petsc4py import PETSc
if PETSc.COMM_WORLD.rank == 0:
PETSc.Vec().create(comm=PETSc.COMM_SELF).view()
它对一个还没设类型的 Vec 调用 view(), 出错信息如下:
It calls view() on a Vec whose type has not been set yet, and the error message is:
$ python test.py
Vec Object: 1 MPI process
type not yet set
Traceback (most recent call last):
File "test.py", line 7, in <module>
PETSc.Vec().create(comm=PETSc.COMM_SELF).view()
File "PETSc/Vec.pyx", line 140, in petsc4py.PETSc.Vec.view
petsc4py.PETSc.Error: error code 56
[0] VecView() at /home/user/software/firedrake-mini-petsc/src/petsc/src/vec/vec/interface/vector.c:715
[0] No support for this operation for this object type
[0] No method view for Vec of type (null)
Python 回溯是完整的, 一眼就能看出问题出在哪一行. 本章后面的 gdb 命令仍以这个 test.py 作为被调试的程序, 只是为了有个可以照着敲的例子; 真正需要 gdb 的是上面列出的那几种”没有回溯”的场景, 本章 tmux-mpi 的调试命令一节给了一次真实的 SIGBUS 会话记录.
The Python traceback is complete and one can see at a glance which line went wrong. The gdb commands later in this chapter still use this test.py as the program being debugged, simply so that there is an example to type along with; what really calls for gdb are the “no traceback” situations listed above, and the section on the debugging commands of tmux-mpi in this chapter gives the record of a real SIGBUS session.
2.3.1. gdb 命令行说明 The gdb command line#
gdb [options] --args executable-file [inferior-arguments ...]
调试 Firedrake 程序时, 被调试的可执行文件是 Python 解释器本身, 脚本只是它的参数, 所以命令里要写成 --args $(which python) test.py.
When debugging a Firedrake program the executable being debugged is the Python interpreter itself and the script is merely one of its arguments, so the command has to be written as --args $(which python) test.py.
2.3.2. 参数 Arguments (options)#
-x file: 从文件中读取 gdb 命令-ex COMMAND: 执行一条 gdb 命令--args exe [exe-args]: 传递参数给 exe--pid <pid>(或-p <pid>): 挂到一个正在运行的进程上. 程序挂住不动时用这个, 不必重启程序
-x file: read gdb commands from a file-ex COMMAND: execute one gdb command--args exe [exe-args]: pass arguments to exe--pid <pid>(or-p <pid>): attach to a running process. Use this when the program hangs, so that it does not have to be restarted
2.3.3. gdb 命令 gdb commands#
bt: 查看函数调用栈run: 运行可执行文件c: 继续运行l: 查看代码p: 打印变量thread apply all bt: 打印所有线程的调用栈, 排查挂起时常用
bt: show the call stackrun: run the executablec: continue runningl: list the codep: print a variablethread apply all bt: print the call stacks of all threads, commonly used when tracking down a hang
2.3.4. 示例 (调试 test.py) Example (debugging test.py)#
$ gdb -ex run --args $(which python3) test.py
Note
这套命令在 Linux 和基于 Intel 芯片的 macOS 上可用. Apple Silicon (arm64) 的 macOS 上 gdb 用不了: 本机实测 Homebrew 的 gdb 16.3 连 Mach-O arm64 可执行文件都识别不出来 (报 not in executable format: file format not recognized), set architecture 的候选列表里也没有 aarch64. 这类机器上请改用系统自带的 lldb, 对应命令见后面 tmux-mpi 的调试命令一节.
Note
These commands work on Linux and on Intel-based macOS. On Apple Silicon (arm64) macOS gdb does not work: measured on this machine, Homebrew’s gdb 16.3 does not even recognize a Mach-O arm64 executable (it reports not in executable format: file format not recognized), and aarch64 is not among the candidates of set architecture either. On such machines use the system’s own lldb instead; the corresponding commands are given in the section on the debugging commands of tmux-mpi below.
2.3.5. gdb 控制命令 gdb control commands#
以下命令可以把 gdb 的调试过程保存到文件, 便于贴进 issue.
输出执行的 gdb 命令
set trace-commands onref https://sourceware.org/gdb/onlinedocs/gdb/Messages_002fWarnings.html关闭分页
set pagination offref https://sourceware.org/gdb/onlinedocs/gdb/Screen-Size.html设置日志文件, 并开启日志
set logging file my.logs,set logging enable onref https://sourceware.org/gdb/download/onlinedocs/gdb/Logging-Output.html
关掉分页是必要的一步: 默认情况下 gdb 每输出一屏就停下来等按键, 写成脚本自动运行时会一直卡在那里.
The commands below save a gdb session to a file, which makes it easy to paste into an issue.
Echo the gdb commands that are executed:
set trace-commands onref https://sourceware.org/gdb/onlinedocs/gdb/Messages_002fWarnings.htmlTurn paging off:
set pagination offref https://sourceware.org/gdb/onlinedocs/gdb/Screen-Size.htmlSet a log file and turn logging on:
set logging file my.logs,set logging enable onref https://sourceware.org/gdb/download/onlinedocs/gdb/Logging-Output.html
Turning paging off is a necessary step: by default gdb stops after every screenful of output and waits for a key press, which would make an automated script hang there forever.
2.3.6. gdb 的 python 插件 The python plugin of gdb#
Python 解释器的调用栈在 gdb 里是一堆 _PyEval_EvalFrame 之类的 C 帧, 看不出对应哪一行 Python 代码. gdb 的 python 插件提供了 py-bt 命令, 可以把它翻译成 Python 层的调用栈.
在 ubuntu 上安装 python3-dbg 后, 文件夹 /usr/share/gdb/auto-load/usr/bin/ 下会有如下插件
In gdb the call stack of the Python interpreter is a pile of C frames such as _PyEval_EvalFrame, from which one cannot tell which line of Python code they correspond to. The python plugin of gdb provides the command py-bt, which translates it into a call stack at the Python level.
On ubuntu, after installing python3-dbg, the following plugins are found in the directory /usr/share/gdb/auto-load/usr/bin/
$ ls /usr/share/gdb/auto-load/usr/bin/
python3.10-dbg-gdb.py python3.10dm-gdb.py python3.10-gdb.py python3.10m-gdb.py
其中定义了用于显示 python 调用栈的命令: py-bt. 下面的文件名以 Ubuntu 22.04 的 Python 3.10 为例, 实际使用时按系统里的版本号替换.
gdb 没有自动加载该插件时, 可以手动加载:
They define the command that shows the python call stack: py-bt. The file names below use Python 3.10 of Ubuntu 22.04 as an example; replace them with the version number found on your own system.
When gdb does not load the plugin automatically, it can be loaded by hand:
source /usr/share/gdb/auto-load/usr/bin/python3.10m-gdb.py
或者把这一行写进 gdb 的初始化文件 $HOME/.config/gdb/gdbinit 或当前目录下的 .gdbinit, 以后每次启动都会自动加载.
Alternatively, put this line into gdb’s initialization file $HOME/.config/gdb/gdbinit, or into a .gdbinit in the current directory, and it will be loaded automatically on every start.
2.3.7. 从文件读取 gdb 命令 Reading gdb commands from a file#
命令多了以后一条条敲很麻烦, 更实用的做法是把它们写进一个文件, 让 gdb 自动执行完再退出 —— 挂到一个卡住的进程上抓一份调用栈时尤其方便.
新建文件 cmd.txt 内容如下
Once there are many commands, typing them one by one is tedious; it is more practical to put them into a file and let gdb run them and then exit — this is particularly convenient for attaching to a hung process and grabbing a call stack.
Create a file cmd.txt with the following content
source /usr/share/gdb/auto-load/usr/bin/python3.10m-gdb.py
set trace-commands on
set pagination off
set logging file my.logs
set logging enable on
py-bt
exit
使用 -x 参数运行 gdb, 它会自动运行上面的命令并退出, 结果留在 my.logs 里.
Running gdb with the -x option makes it run the commands above and exit, leaving the result in my.logs.
gdb -x cmd.txt -p <pid>
2.4. 并行程序调试 Debugging parallel programs#
并行下的问题大致分两类, 处理方式完全不同:
某个进程崩溃或报错. 各进程的输出会搅在一起, 而且往往只有一个进程真的出错. 先设法弄清是哪个 rank 出的问题, 再只对那个 rank 开调试器.
程序挂住不动. 最常见的原因是集合操作只被一部分进程调用了 —— 例如把
assemble、solve、norm或dat.data写在if rank == 0:里面, 其余进程走不到这一句, 于是所有进程一起等下去. 本项目用mpiexec -n 2实测过, 程序会永久挂起, 既不报错也不退出, 详见 并行常见陷阱 Common parallel pitfalls. 这类问题没有任何报错信息, 只能挂上调试器看各个进程分别停在哪里.
不管哪一类, 关键都是给每个进程一个独立的终端, 否则多个调试器的输入输出会互相干扰.
Problems in parallel fall roughly into two classes, which are handled in completely different ways:
One process crashes or reports an error. The output of the processes is interleaved, and often only one process really failed. First find out which rank the problem comes from, then open a debugger for that rank only.
The program hangs. The most common cause is a collective operation being called by only some of the processes — for example writing
assemble,solve,normordat.datainsideif rank == 0:, so that the other processes never reach this statement and all of them wait forever. This has been measured in this project withmpiexec -n 2: the program hangs permanently, neither erroring out nor exiting, see 并行常见陷阱 Common parallel pitfalls. Problems of this kind come with no error message at all, and the only way is to attach a debugger and see where each process is stopped.
In either case the key is to give every process a terminal of its own, otherwise the input and output of the several debuggers interfere with each other.
2.4.1. PETSc 的参数 The PETSc option -start_in_debugger#
PETSc 自带一个开关: 加上 -start_in_debugger, 它会在初始化阶段给每个进程各开一个终端窗口, 窗口里是已经挂好的调试器.
PETSc comes with a switch of its own: with -start_in_debugger it opens, during initialization, one terminal window per process, each holding a debugger already attached.
mpiexec -n 3 $(which python) test.py -start_in_debugger
选项要写在脚本名之后: 这样它们才会留在 sys.argv 里, 由 PETSc 在初始化时读走 (Firedrake 一被 import 就会做这件事).
默认行为和平台有关 (以下为 PETSc 3.25 源码中的默认值):
调试器在 PETSc 配置阶段就定下来了: macOS 上默认是
lldb, 其他系统默认是gdb;终端在 macOS 上是
Terminal.app, 在其他系统上是xterm.
几个配套选项, 进程一多就很有用:
选项 |
作用 |
|---|---|
|
只给指定的 rank 开调试器, 默认是所有进程 |
|
不另开窗口, 就在当前终端里挂调试器 |
|
不开调试器, 只打印”如何手工挂上去”的提示, 然后停下来等 |
|
换一个终端程序, 例如 |
|
等若干秒再挂上去, 适合终端启动慢的场合 |
其中 -stop_for_debugger 在只有 ssh、没有 X 转发的机器上尤其有用: 开不出窗口时, 让它打印出进程号, 再自己 gdb -p <pid> 挂上去.
参考:
The options have to be written after the script name: only then do they stay in sys.argv and get picked up by PETSc during initialization (which happens as soon as Firedrake is imported).
The default behavior depends on the platform (the defaults below are those in the PETSc 3.25 sources):
the debugger is fixed at PETSc configure time: on macOS the default is
lldb, on other systems it isgdb;the terminal is
Terminal.appon macOS andxtermelsewhere.
A few companion options, which become useful as soon as there are many processes:
Option |
Effect |
|---|---|
|
open a debugger only for the given ranks; the default is all processes |
|
do not open a separate window, attach the debugger in the current terminal |
|
do not open a debugger, only print a hint on how to attach one by hand, then stop and wait |
|
use a different terminal program, for example |
|
wait a number of seconds before attaching, suitable when the terminal starts slowly |
Among these, -stop_for_debugger is particularly useful on a machine reachable only over ssh with no X forwarding: when no window can be opened, let it print the process id and then attach with gdb -p <pid> oneself.
References:
Tip
修改 xterm 窗口显示效果 (终端是 xterm 时才有意义, 即 Linux)
参考: http://www.futurile.net/2016/06/14/xterm-setup-and-truetype-font-configuration/
$ cat ~/.Xdefaults
xterm*faceName: Monospace
xterm*faceSize: 12
xterm*foreground: rgb:a8/a8/a8
xterm*background: rgb:00/00/00
Tip
Changing the appearance of the xterm window (only meaningful when the terminal is xterm, so on Linux)
Reference: http://www.futurile.net/2016/06/14/xterm-setup-and-truetype-font-configuration/
$ cat ~/.Xdefaults
xterm*faceName: Monospace
xterm*faceSize: 12
xterm*foreground: rgb:a8/a8/a8
xterm*background: rgb:00/00/00
2.4.2. 工具 The tool tmux-mpi#
-start_in_debugger 要能开出图形窗口, 在只有 ssh 终端的集群上通常用不了. tmux-mpi 换了个思路: 它给每个 MPI 进程开一个 tmux 窗口, 每个窗口里跑一个调试器, 全程在字符终端里完成, 不需要 X 转发, 断线之后重新 attach 还能接着调.
参考:
Firedrake wiki 上的说明: firedrakeproject/firedrake
Tips of Firedrake Wiki: firedrakeproject/firedrake
tmux-mpi 配置与高级用法: wrs20/tmux-mpi
-start_in_debugger needs to be able to open graphical windows, which usually rules it out on a cluster with nothing but an ssh terminal. tmux-mpi takes a different route: it opens one tmux window per MPI process, each running a debugger, entirely inside a character terminal, with no X forwarding required, and after a dropped connection one can attach again and carry on debugging.
References:
The description on the Firedrake wiki: firedrakeproject/firedrake
Tips of Firedrake Wiki: firedrakeproject/firedrake
Configuration and advanced usage of tmux-mpi: wrs20/tmux-mpi
2.4.2.1. 安装 tmux-mpi Installing tmux-mpi#
tmux-mpi 依赖 tmux 和 dtach 两个外部程序, 需要先装好.
tmux-mpi depends on the two external programs tmux and dtach, which have to be installed first.
安装 tmuxInstalling tmux
Ubuntu:
sudo apt-get install tmux
macOS:
brew install tmux
安装 dtach (tmux-mpi 依赖)Installing dtach (a dependency of tmux-mpi)
dtach 一般不在发行版仓库里, 需要自己编译, 然后把二进制文件拷到某个在
PATH中的目录, 如$HOME/bin.dtach is usually not in the distribution repositories and has to be compiled by hand; the binary is then copied into a directory on thePATH, such as$HOME/bin.git clone https://github.com/crigler/dtach cd dtach ./configure make mkdir -p $HOME/bin cp dtach $HOME/bin export PATH=$PATH:$HOME/bin
运行
which dtach确认安装是否成功. 最后那行export只对当前终端有效, 要长期生效需写进~/.bashrc或~/.zshrc.Runwhich dtachto confirm that the installation succeeded. The lastexportline only affects the current terminal; to make it permanent write it into~/.bashrcor~/.zshrc.安装 tmux-mpiInstalling tmux-mpi
在 Firedrake 的虚拟环境里用 pip 安装Install it with pip inside Firedrake’s virtual environment
pip install --upgrade --no-cache-dir git+https://github.com/wrs20/tmux-mpi@master
也可以用
pipx install git+https://github.com/wrs20/tmux-mpi@master装到全局, 不污染虚拟环境. 装好后which tmux-mpi应当能找到它.It can also be installed globally withpipx install git+https://github.com/wrs20/tmux-mpi@master, which leaves the virtual environment untouched. After installationwhich tmux-mpishould find it.
2.4.2.2. 调试命令 Debugging commands#
Linux 和基于 Intel 芯片的 macOS 系统可以使用 gdb 作为调试器; 在使用 Apple Silicon (arm64) 的 macOS 系统上, 需改用 lldb.
On Linux and on Intel-based macOS gdb can be used as the debugger; on macOS with Apple Silicon (arm64) one has to use lldb instead.
启动调试器Start the debugger
tmux-mpi 3 gdb -ex run --args $(which python) test.py
用
lldb时对应写成Withlldbthe corresponding form istmux-mpi 3 lldb -o run -o continue -- $(which python) test.py
这里比 gdb 多一条
-o continue: 本机实测lldb -o run会先停在dyld_start(动态加载器的入口), 要再continue一次程序才真正跑起来.There is one more-o continuehere than with gdb: measured on this machine,lldb -o runfirst stops atdyld_start(the entry point of the dynamic loader), and only after one morecontinuedoes the program really start running.Attach 到相应的伪终端, 每个进程一个窗口. (这里是 tmux 的一个 session, 有多个 window)Attach to the corresponding pseudo-terminals, one window per process. (This is a single tmux session with several windows.)
tmux attach -t tmux-mpi
tmux 的窗口切换:
Ctrl-b之后按n/p切到下一个 / 上一个窗口, 按数字键直接跳到对应编号的进程, 按d从 session 里退出 (进程还在跑, 之后可以再 attach 回来).Switching windows in tmux: afterCtrl-b, pressn/pto move to the next / previous window, press a digit to jump straight to the process with that number, and pressdto detach from the session (the processes keep running and can be attached to again later).使用 gdb 调试命令调试, 例如查看栈回溯 (
bt)、反汇编 (disassemble)、寄存器和内存映射. 下面是一次实际调试SIGBUS错误的会话记录: 程序在 MPICH 的memcpy里收到总线错误, 于是先用bt看是从哪里调用过来的, 再打印相关寄存器和info proc mapping, 判断拷贝的目标地址是不是落到了共享内存段 (/dev/shm/mpich_shar_tmp...) 的边界之外.Debug with the gdb commands, for example looking at the backtrace (bt), the disassembly (disassemble), the registers and the memory mappings. Below is the record of a real session debugging aSIGBUSerror: the program received a bus error inside MPICH’smemcpy, sobtwas used first to see where the call came from, and then the relevant registers andinfo proc mappingwere printed in order to decide whether the destination address of the copy had fallen outside the bounds of the shared memory segment (/dev/shm/mpich_shar_tmp...).Thread 1 "python" received signal SIGBUS, Bus error. __memmove_evex_unaligned_erms () at ../sysdeps/x86_64/multiarch/memmove-vec-unaligned-erms.S:708 708 ../sysdeps/x86_64/multiarch/memmove-vec-unaligned-erms.S: No such file or directory. (gdb) bt #0 __memmove_evex_unaligned_erms () at ../sysdeps/x86_64/multiarch/memmove-vec-unaligned-erms.S:708 #1 0x00007feb09870bfb in memcpy (__len=47424, __src=<optimized out>, __dest=<optimized out>) at /usr/include/x86_64-linux-gnu/bits/string_fortified.h:29 #2 MPID_nem_mpich_sendv_header (again=<synthetic pointer>, vc=0x55a9d0d49bb0, n_iov=<synthetic pointer>, iov=<synthetic pointer>) at ./src/mpid/ch3/channels/nemesis/include/mpid_nem_inline.h:396 #3 MPIDI_CH3_iSendv (vc=0x55a9d0d49bb0, sreq=<optimized out>, sreq@entry=0x55a9d3b5e430, iov=iov@entry=0x7ffe5ee5e0f0, n_iov=n_iov@entry=2) at src/mpid/ch3/channels/nemesis/src/ch3_isendv.c:69 #4 0x00007feb0985c7dc in MPIDI_CH3_EagerContigIsend ( sreq_p=sreq_p@entry=0x7ffe5ee5e268, reqtype=reqtype@entry=MPIDI_CH3_PKT_EAGER_SEND, buf=<optimized out>, data_sz=<optimized out>, rank=rank@entry=2, tag=<optimized out>, comm=0x55a9d327aab0, context_offset=0) at src/mpid/ch3/src/ch3u_eager.c:545 ... (gdb) disassemble Dump of assembler code for function __memmove_evex_unaligned_erms: 0x00007feb0c349f40 <+0>: endbr64 0x00007feb0c349f44 <+4>: mov %rdi,%rax 0x00007feb0c349f47 <+7>: cmp $0x20,%rdx 0x00007feb0c349f4b <+11>: jb 0x7feb0c349f80 <__memmove_evex_unaligned_erms+64> 0x00007feb0c349f4d <+13>: vmovdqu64 (%rsi),%ymm16 0x00007feb0c349f53 <+19>: cmp $0x40,%rdx 0x00007feb0c349f57 <+23>: ja 0x7feb0c34a000 <__memmove_evex_unaligned_erms+192> ... 0x00007feb0c34a286 <+838>: add %rdi,%rsi 0x00007feb0c34a289 <+841>: sub %rdi,%rcx => 0x00007feb0c34a28c <+844>: rep movsb %ds:(%rsi),%es:(%rdi) 0x00007feb0c34a28e <+846>: vmovdqu64 %ymm16,(%r8) 0x00007feb0c34a294 <+852>: vmovdqu64 %ymm17,0x20(%r8) ... --Type <RET> for more, q to quit, c to continue without paging--q Quit (gdb) p/x $rsi $13 = 0x7fe9f43c93a0 (gdb) p/x $rdi $14 = 0x7feb016a4000 (gdb) p/x $rcx $15 = 0x25b0 (gdb) p $rcx $16 = 9648 (gdb) p/x $rdi + $rcx $17 = 0x7feb016a65b0 (gdb) p/x $rsi $13 = 0x7fe9f43c93a0 (gdb) p/x $rdi $14 = 0x7feb016a4000 (gdb) info proc mapping ... 0x7feafe633000 0x7feafe634000 0x1000 0x6000 r--p /usr/lib/python3.10/lib-dynload/_bz2.cpython-310-x86_64-linux-gnu.so 0x7feafe634000 0x7feafe635000 0x1000 0x7000 rw-p /usr/lib/python3.10/lib-dynload/_bz2.cpython-310-x86_64-linux-gnu.so 0x7feafe635000 0x7feafebf7000 0x5c2000 0x0 rw-p 0x7feafebf7000 0x7feb03afc000 0x4f05000 0x0 rw-s /dev/shm/mpich_shar_tmpMCilHL (deleted) 0x7feb03afc000 0x7feb03bfc000 0x100000 0x0 rw-p ...
lldb里对应的命令是bt(相同)、thread backtrace all、disassemble、register read和image list.The corresponding commands inlldbarebt(the same),thread backtrace all,disassemble,register readandimage list.
2.5. 编译 .pyx 文件 Compiling .pyx files#
2.5.1. Firedrake 中的 .pyx 文件 The .pyx files in Firedrake#
Firedrake 有一部分底层代码是用 Cython 写的, 放在 firedrake/cython/ 目录下, 编译成扩展模块 (.so) 之后才能被 import:
Part of Firedrake’s low-level code is written in Cython and lives in the directory firedrake/cython/; it can only be imported once it has been compiled into extension modules (.so):
dmcommon.pyx DM (DMPlex) 相关的公共工具函数
extrusion_numbering.pyx 拉伸网格的自由度编号
hdf5interface.pyx HDF5 底层接口
mgimpl.pyx 多重网格的底层编号
patchimpl.pyx PCPatch 的回调
spatialindex.pyx 空间索引 (点定位)
supermeshimpl.pyx supermesh 投影
改动这些文件之后必须重新编译, 否则改的是 .pyx, 跑的仍然是旧的 .so. 这件事只对源码 (editable) 安装有意义: 用 wheel 装的 Firedrake 目录里虽然也带着 .pyx, 但那只是随包附带的源文件, 改了不会生效.
在 Firedrake 源码树的根目录 (有 setup.py 的那一层) 执行下面的命令就地重新编译:
The annotations in the listing give, in order: common utilities for DM (DMPlex), the numbering of degrees of freedom on extruded meshes, the low-level HDF5 interface, the low-level numbering for multigrid, the callbacks of PCPatch, the spatial index (point location) and supermesh projection.
After changing these files they must be recompiled, otherwise what was changed is the .pyx while what runs is still the old .so. This only makes sense for a source (editable) installation: a Firedrake installed from a wheel also carries the .pyx files in its directory, but those are merely the source files shipped with the package, and changing them has no effect.
Run the following command in the root of the Firedrake source tree (the level that contains setup.py) to recompile in place:
python setup.py build_ext --inplace
也可以按官方文档的方式重装一次, 效果相同:
Reinstalling once in the way the official documentation describes has the same effect:
pip install --no-build-isolation --editable .
Note
本机是非 editable 安装, 没有 Firedrake 源码树, 上面两条命令未在本机实测. 已经确认的是: Firedrake 仓库根目录仍有 setup.py (setuptools + Cython 构建), 官方安装文档给出的开发安装方式是 pip install --no-build-isolation --editable ...; 上面的 .pyx 文件清单来自 Firedrake 2026.4.1 的安装目录.
Note
This machine has a non-editable installation and no Firedrake source tree, so the two commands above have not been tested on this machine. What has been confirmed is: the root of the Firedrake repository still contains a setup.py (a setuptools + Cython build), the development installation described in the official installation documentation is pip install --no-build-isolation --editable ..., and the list of .pyx files above comes from the installation directory of Firedrake 2026.4.1.
2.6. 排查清单 Checklist#
卡住的时候可以按下面这份通用清单逐条过一遍.
先把报错读完. Python 回溯看最后几行; PETSc 的错误块看
[0]PETSC ERROR:的第一条, 后面几条只是调用栈.变量是否初始化. 尤其是循环里反复使用的
Function, 每一步之前是否需要清零.下标是否错位.
n还是n+1还是n-1, 各自对应哪个时刻的值, 算一遍对上再往下写.把规模缩小. 网格调到 4x4、时间步只跑两步, 出错更快, 输出也读得完.
先串行, 后并行. 串行跑不通的程序不要拿去并行; 并行结果对不对, 用集合量 (
assemble、norm、errornorm) 和串行结果逐位对照.确认参数真的生效了. 求解器参数拼错不会报错, 只会被忽略 (本机实测把
ksp_type误写成ksp_typ, 程序照常跑完). 用ksp_view看实际用的是什么求解器, 加命令行参数-options_left列出所有没被用上的选项.
When stuck, go through this general checklist item by item.
Read the error message to the end first. For a Python traceback look at the last few lines; for a PETSc error block look at the first
[0]PETSC ERROR:line, the following ones are only the call stack.Are the variables initialized? Especially a
Functionreused inside a loop: does it have to be zeroed before every step?Are the indices off by one? Is it
n,n+1orn-1, and which time level does each of them stand for? Work it out once and check it before writing on.Shrink the problem. Turn the mesh down to 4x4 and run only two time steps: the failure comes sooner and the output can be read in full.
Serial first, parallel afterwards. Do not take a program that fails in serial into parallel; to check whether a parallel result is right, compare collective quantities (
assemble,norm,errornorm) digit by digit against the serial result.Confirm that the options really took effect. A misspelled solver parameter raises no error and is simply ignored (measured on this machine:
ksp_typemistyped asksp_typand the program ran to completion as usual). Useksp_viewto see which solver is actually used, and add the command-line option-options_leftto list all the options that were not used.