2. 调试 Debugging#

概要

调试 Firedrake 程序: 常见报错及其成因, 用 pdb/gdb 调试 Python 代码与生成的 C 代码, 用 -start_in_debugger 和 tmux-mpi 做并行调试, 以及编译 .pyx 文件.

Overview

Debugging Firedrake programs: common errors and their causes, debugging Python and generated C code with pdb/gdb, parallel debugging with -start_in_debugger and tmux-mpi, and compiling .pyx files.

2.1. 常见问题 Common problems#

本节按报错信息组织. 遇到问题时可以直接拿报错里的关键词 (例如 DIVERGED_LINEAR_SOLVE) 在本页搜索.

This section is organized by error message. When a problem shows up, one can search this page directly for a keyword taken from the error (for example DIVERGED_LINEAR_SOLVE).

2.1.1. 线性求解不收敛The linear solve does not converge: DIVERGED_LINEAR_SOLVE#

求解失败时 Firedrake 抛出的异常大致长这样:

When the solve fails, the exception raised by Firedrake looks roughly like this:

File "/home/user/firedrake/src/firedrake/firedrake/adjoint/solving.py", line 50, in wrapper
    output = solve(*args, **kwargs)
  File "/home/user/firedrake/src/firedrake/firedrake/solving.py", line 129, in solve
    _solve_varproblem(*args, **kwargs)
  File "/home/user/firedrake/src/firedrake/firedrake/solving.py", line 161, in _solve_varproblem
    solver.solve()
  File "/home/user/firedrake/src/firedrake/firedrake/adjoint/variational_solver.py", line 75, in wrapper
    out = solve(self, **kwargs)
  File "/home/user/firedrake/src/firedrake/firedrake/variational_solver.py", line 278, in solve
    solving_utils.check_snes_convergence(self.snes)
  File "/home/user/firedrake/src/firedrake/firedrake/solving_utils.py", line 139, in check_snes_convergence
    raise ConvergenceError(r"""Nonlinear solve failed to converge after %d nonlinear iterations.
firedrake.exceptions.ConvergenceError: Nonlinear solve failed to converge after 0 nonlinear iterations.
Reason:
   DIVERGED_LINEAR_SOLVE

这条信息只说明”线性求解没有成功”, 并没有说为什么. 常见成因有:

  1. 方程没有定好. 最常见的是边界条件写错或写漏, 使离散系统奇异或严重病态. 先逐条核对 DirichletBC 的边界编号和取值.

  2. 系统本身是奇异的. 例如纯 Neumann 问题, 解只差一个常数. 这类问题要给出零空间 (nullspace), 或者用 Lagrange 乘子把常数固定下来.

  3. 迭代次数用完了. 迭代法不收敛未必是方程错了, 也可能只是预条件子不合适, 或者 ksp_max_it 给得太小.

  4. 外部库报错. 用直接法 (LU) 时, 错误常常来自 PETSc 调用的 MUMPS 等外部库, 见下面两个例子.

排查步骤. Firedrake 抛出的异常只到”线性求解失败”为止, 真正的原因要到 PETSc 的 KSP / PC 层去问 (取用方法见 PETSc 一章的 KSP 一节).

This message only says that “the linear solve did not succeed”, not why. The common causes are:

  1. The equation is not set up correctly. Most often a boundary condition is wrong or missing, which makes the discrete system singular or severely ill-conditioned. Check the boundary ids and the values of every DirichletBC one by one first.

  2. The system itself is singular. For example a pure Neumann problem, whose solution is determined only up to a constant. Such problems need a nullspace (nullspace), or a Lagrange multiplier to pin the constant down.

  3. The iterations ran out. A non-converging iterative method does not necessarily mean the equation is wrong; the preconditioner may simply be unsuitable, or ksp_max_it may be too small.

  4. An external library reported an error. With a direct method (LU) the error often comes from an external library called by PETSc such as MUMPS, see the two examples below.

Checking steps. The exception raised by Firedrake stops at “the linear solve failed”; the real cause has to be asked for at PETSc’s KSP / PC level (for how to get hold of them see the section KSP in the PETSc chapter).

  1. 让 PETSc 自己报告残差历史和收敛原因:Let PETSc itself report the residual history and the convergence reason:

    solver_parameters = {
        # ... 原有的参数
        'ksp_monitor': None,
        'ksp_converged_reason': None,
    }
    

    输出形如 (下面是把 ksp_max_it 设成 2 故意制造的不收敛):The output looks like this (below, the non-convergence is produced on purpose by setting ksp_max_it to 2):

        Residual norms for firedrake_1_ solve.
        0 KSP Residual norm 6.091695533358e-02
        1 KSP Residual norm 2.802120913580e-03
        2 KSP Residual norm 1.265467409231e-03
        Linear firedrake_1_ solve did not converge due to DIVERGED_ITS iterations 2
    

    DIVERGED_ITS 才是真正的原因 (迭代次数用完), 而 Firedrake 的异常里只有 DIVERGED_LINEAR_SOLVE.DIVERGED_ITS is the real reason (the iterations ran out), while Firedrake’s exception only carries DIVERGED_LINEAR_SOLVE.

  2. 加 'ksp_error_if_not_converged': None, 让 PETSc 在失败的地方直接抛错, 异常里会带上 C 端的调用栈:Add 'ksp_error_if_not_converged': None to make PETSc raise the error where it fails; the exception then carries the call stack of the C side:

    petsc4py.PETSc.Error: error code 91
    [0] SNESSolve() at .../src/snes/interface/snes.c:4875
    [0] SNESSolve_KSPONLY() at .../src/snes/impls/ksponly/ksponly.c:49
    [0] KSPSolve() at .../src/ksp/ksp/interface/itfunc.c:1092
    [0] KSPSolve_Private() at .../src/ksp/ksp/interface/itfunc.c:999
    [0] KSPSolve() has not converged, reason DIVERGED_ITS
    

    错误码 91 就是 PETSC_ERR_NOT_CONVERGED, 对照表见后面的 PetscErrorCode 一节.Error code 91 is PETSC_ERR_NOT_CONVERGED; the table is given in the PetscErrorCode section below.

  3. 加 'ksp_view': None 确认实际用的求解器和预条件子是不是你以为的那个. 注意 Firedrake 的默认求解器是直接法 (preonly + lu), 不显式设置 ksp_type / pc_type 时并不会真的做迭代.Add 'ksp_view': None to confirm that the solver and the preconditioner actually used are the ones you think they are. Note that Firedrake’s default solver is a direct method (preonly + lu), so unless ksp_type / pc_type are set explicitly no iteration is performed at all.

下面是两个来自 MUMPS 的实际例子.

Below are two real examples coming from MUMPS.

示例 1: MUMPS 工作空间不足 (INFOG(1)=-9)Example 1: MUMPS runs out of workspace (INFOG(1)=-9)

用直接法时, 错误往往来自 PETSc 调用的外部库. 下面这段是 MUMPS 在数值分解阶段报的错, 关键是 INFOG(1) 和 INFO(2) 两个数字, 拿它们去查 MUMPS 手册即可.

参考资料:

  1. MUMPS 用户手册: https://graal.ens-lyon.fr/MUMPS/index.php?page=doc

  2. PETSc 的 MATSOLVERMUMPS 手册页: https://petsc.org/main/manualpages/Mat/MATSOLVERMUMPS/

With a direct method the error often comes from an external library called by PETSc. The block below is an error reported by MUMPS in the numerical factorization phase; what matters are the two numbers INFOG(1) and INFO(2), which are then looked up in the MUMPS manual.

References:

  1. The MUMPS user’s guide: https://graal.ens-lyon.fr/MUMPS/index.php?page=doc

  2. PETSc’s MATSOLVERMUMPS manual page: https://petsc.org/main/manualpages/Mat/MATSOLVERMUMPS/

[63]PETSC ERROR: --------------------- Error Message --------------------------------------------------------------
[63]PETSC ERROR: Error in external library
[63]PETSC ERROR: Error reported by MUMPS in numerical factorization phase: INFOG(1)=-9, INFO(2)=26
[63]PETSC ERROR: See https://petsc.org/release/faq/ for trouble shooting.
[63]PETSC ERROR: Petsc Development GIT revision: v3.4.2-38777-g979cc68729  GIT Date: 2022-06-22 21:18:23 +0100
[63]PETSC ERROR: ../hsolver/hsolver.py on a default named myhost by user Tue Nov  8 16:03:18 2022
[63]PETSC ERROR: Configure options PETSC_DIR=/home/user/opt/firedrake-env/firedrake-complex-int64/src/petsc PETSC_ARCH=default --download-ptscotch --with-zlib --download-hwloc --with-c2html=0 --download-eigen="/home/user/opt/firedrake-env/firedrake-complex-int64/src/eigen-3.3.3.tgz " --download-mpich --download-hdf5 --with-fortran-bindings=0 --with-64-bit-indices --download-bison --with-cxx-dialect=C++11 --download-metis --download-openblas --download-openblas-make-options="'USE_THREAD=0 USE_LOCKING=1 USE_OPENMP=0'" --download-pastix --download-mumps --with-shared-libraries=1 --with-scalar-type=complex --download-cmake --download-scalapack --with-debugging=0 --download-netcdf --download-superlu_dist --download-suitesparse --download-pnetcdf
[63]PETSC ERROR: #1 MatFactorNumeric_MUMPS() at /home/user/opt/firedrake-env/firedrake-complex-int64/src/petsc/src/mat/impls/aij/mpi/mumps/mumps.c:1664
[63]PETSC ERROR: #2 MatLUFactorNumeric() at /home/user/opt/firedrake-env/firedrake-complex-int64/src/petsc/src/mat/interface/matrix.c:3177
[63]PETSC ERROR: #3 PCSetUp_LU() at /home/user/opt/firedrake-env/firedrake-complex-int64/src/petsc/src/ksp/pc/impls/factor/lu/lu.c:135
[63]PETSC ERROR: #4 PCSetUp() at /home/user/opt/firedrake-env/firedrake-complex-int64/src/petsc/src/ksp/pc/interface/precon.c:993
[63]PETSC ERROR: #5 KSPSetUp() at /home/user/opt/firedrake-env/firedrake-complex-int64/src/petsc/src/ksp/ksp/interface/itfunc.c:407
[63]PETSC ERROR: #6 KSPSolve_Private() at /home/user/opt/firedrake-env/firedrake-complex-int64/src/petsc/src/ksp/ksp/interface/itfunc.c:843
[63]PETSC ERROR: #7 KSPSolve() at /home/user/opt/firedrake-env/firedrake-complex-int64/src/petsc/src/ksp/ksp/interface/itfunc.c:1078

MUMPS 手册中关于 INFOG(1) = -9 和 ICNTL(14) 的说明:

What the MUMPS manual says about INFOG(1) = -9 and ICNTL(14):

The main internal real/complex workarray S is too small. If INFO(2) is positive, then the number
of entries that are missing in S at the moment when the error is raised is available in INFO(2).
If INFO(2) is negative, then its absolute value should be multiplied by 1 million. If an error –9
occurs, the user should increase the value of ICNTL(14) before calling the factorization (JOB=2)
again, except if LWK_USER is provided LWK_USER should be increased.
ICNTL(14) corresponds to the percentage increase in the estimated working space.
Phase: accessed by the host both during the analysis and the factorization phases.
Default value: between 20 and 35 (which corresponds to at most 35 % increase) and depends on
               the number of MPI processes. It is set to 5 % with SYM=1 and one MPI process.
Related parameters: ICNTL(23)
Remarks: When significant extra fill-in is caused by numerical pivoting, increasing ICNTL(14)
         may help

也就是说 -9 表示 MUMPS 内部的实数工作数组不够大, 办法是把 ICNTL(14) (估计工作空间的放大百分比) 调大之后重新分解.

MUMPS 的参数在 PETSc 里通过 -mat_mumps_icntl_<num> 设置, 写进 Firedrake 的 solver_parameters 时要放在顶层:

In other words, -9 means that the internal real workarray of MUMPS is too small, and the remedy is to increase ICNTL(14) (the percentage increase of the estimated working space) and factorize again.

The MUMPS parameters are set in PETSc through -mat_mumps_icntl_<num>; when written into Firedrake’s solver_parameters they have to be placed at the top level:

solver_parameters = {
    'ksp_type': 'preonly',
    'pc_type': 'lu',
    'pc_factor_mat_solver_type': 'mumps',
    'mat_mumps_icntl_14': 200,   # 工作空间放大 200%
}

Tip

求解器参数拼错了不会报错, 只会被悄悄忽略. 想确认某个 MUMPS 参数真的生效了, 加上 'mat_mumps_icntl_4': 2 让 MUMPS 把它实际使用的 ICNTL 值打印出来:

 ICNTL(14) Percentage of memory relaxation      =              20

本机实测 (Firedrake 2026.4.1, MUMPS 5.8.2): 顶层的 mat_mumps_icntl_14 能改动这一行的数值, 而写成 pc_factor_mat_mumps_icntl_14 时打印出来的仍是 MUMPS 自己的默认值 20.

Tip

A misspelled solver parameter raises no error and is silently ignored. To confirm that a MUMPS parameter really took effect, add 'mat_mumps_icntl_4': 2 so that MUMPS prints the ICNTL values it actually uses:

 ICNTL(14) Percentage of memory relaxation      =              20

Measured on this machine (Firedrake 2026.4.1, MUMPS 5.8.2): mat_mumps_icntl_14 at the top level does change the number on this line, whereas written as pc_factor_mat_mumps_icntl_14 the printed value is still MUMPS’ own default of 20.

其他有用的选项Other useful options

ksp_view、ksp_monitor、ksp_converged_reason、ksp_error_if_not_converged.ksp_view, ksp_monitor, ksp_converged_reason, ksp_error_if_not_converged.

示例 2: MUMPS 遇到零主元 (INFOG(1)=-10)Example 2: MUMPS hits a zero pivot (INFOG(1)=-10)

下面这条是 petsc4py 直接抛出来的 PETSc 错误, 错误码 76 即 PETSC_ERR_LIB, 表示 PETSc 调用的外部库出了错:

The following is a PETSc error raised directly by petsc4py; error code 76 is PETSC_ERR_LIB, meaning that an external library called by PETSc failed:

petsc4py.PETSc.Error: error code 76
[0] SNESSolve() at /home/user/firedrake/real-int32-mkl-debug/src/petsc/src/snes/interface/snes.c:4693
[0] SNESSolve_KSPONLY() at /home/user/firedrake/real-int32-mkl-debug/src/petsc/src/snes/impls/ksponly/ksponly.c:48
[0] KSPSolve() at /home/user/firedrake/real-int32-mkl-debug/src/petsc/src/ksp/ksp/interface/itfunc.c:1070
[0] KSPSolve_Private() at /home/user/firedrake/real-int32-mkl-debug/src/petsc/src/ksp/ksp/interface/itfunc.c:824
[0] KSPSetUp() at /home/user/firedrake/real-int32-mkl-debug/src/petsc/src/ksp/ksp/interface/itfunc.c:405
[0] PCSetUp() at /home/user/firedrake/real-int32-mkl-debug/src/petsc/src/ksp/pc/interface/precon.c:994
[0] PCSetUp_LU() at /home/user/firedrake/real-int32-mkl-debug/src/petsc/src/ksp/pc/impls/factor/lu/lu.c:120
[0] MatLUFactorNumeric() at /home/user/firedrake/real-int32-mkl-debug/src/petsc/src/mat/interface/matrix.c:3215
[0] MatFactorNumeric_MUMPS() at /home/user/firedrake/real-int32-mkl-debug/src/petsc/src/mat/impls/aij/mpi/mumps/mumps.c:1683
[0] Error in external library
[0] Error reported by MUMPS in numerical factorization phase: INFOG(1)=-10, INFO(2)=4464

MUMPS 手册对 INFOG(1)=-10 的说明:

What the MUMPS manual says about INFOG(1)=-10:

INFOG(1)=-10: Numerically singular matrix. INFO(2) holds the number of eliminated pivots.

也就是矩阵在数值上是奇异的. 打开 MUMPS 的零主元检测可以让分解不再直接失败:

So the matrix is numerically singular. Turning on MUMPS’ zero pivot detection keeps the factorization from failing outright:

solver_parameters = {
    'ksp_type': 'preonly',
    'pc_type': 'lu',
    'pc_factor_mat_solver_type': 'mumps',
    'mat_mumps_icntl_24': 1,       # 检测零主元
    # 'mat_mumps_cntl_3': 1e-12,   # 判定零主元的阈值
}
ICNTL(24) controls the detection of ``null pivot rows''.
    Phase: accessed by the host during the factorization phase
    Possible variables/arrays involved: PIVNUL LIST
    Possible values :
        0: Nothing done. A null pivot row will result in error INFO(1)=-10.
        1: Null pivot row detection.
    Other values are treated as 0.
    Default value: 0 (no null pivot row detection)

Warning

ICNTL(24) 只是让分解阶段不因为零主元而报错, 它并不替你处理奇异性: 解在零空间方向上的分量是任意的.

如果矩阵奇异是问题本身决定的 (例如纯 Neumann 问题), 正确的做法是给 solve 传 nullspace. 用直接法时两件事都要做: nullspace 只负责把零空间分量投影掉, 不参与 LU 分解, 而 ICNTL(24) 才让分解本身能走完.

Warning

ICNTL(24) only keeps the factorization phase from erroring out on a zero pivot; it does not deal with the singularity for you: the component of the solution along the nullspace is arbitrary.

If the singularity is inherent to the problem (a pure Neumann problem, say), the right thing to do is to pass a nullspace to solve. With a direct method both are needed: nullspace only projects the nullspace component out and takes no part in the LU factorization, while ICNTL(24) is what lets the factorization itself run to completion.

2.1.2. 高阶网格上对法向求导: ReferenceGrad 报错 Differentiating along the normal on a higher-order mesh: the ReferenceGrad error#

在比较老的 Firedrake/UFL 上, 对高阶网格 (等参元) 的单元边界外法向求导会报下面的错:

On older versions of Firedrake/UFL, differentiating with respect to the outward normal on the facets of a higher-order (isoparametric) mesh raised the following error:

ufl.log.UFLException: Currently no support for ReferenceGrad in CoefficientDerivative.

这条限制来自 UFL: 对系数求导时, 它不支持表达式里出现参考单元上的梯度 ReferenceGrad. 高阶网格上的 FacetNormal 依赖坐标场, 求导时就可能落到这条分支上.

下面的单元把两种写法各在仿射网格和高阶网格上跑一遍, 用来检查当前版本还有没有这个限制: get_s12 先对 u 求三次导再和法向缩并, get_s12_v2 则多了一次”对含法向的表达式再求导”.

The restriction comes from UFL: when differentiating with respect to a coefficient it does not support the gradient on the reference cell, ReferenceGrad, appearing in the expression. On a higher-order mesh FacetNormal depends on the coordinate field, so differentiating it may end up on this branch.

The cell below runs both formulations on an affine mesh and on a higher-order mesh, so as to check whether the current version still has this restriction: get_s12 takes three derivatives of u and then contracts them with the normal, while get_s12_v2 additionally differentiates an expression that already contains the normal.

Note

本机实测 (Firedrake 2026.4.1, UFL 2025.3.0): 这四种组合现在全部通过, 单元输出四行 TEST OK!, 这条报错在这条路径上已经不再出现.

UFL 里这条限制本身还在, 只是异常类型变了: 抛出点在 ufl/algorithms/apply_derivatives.py 的 ReferenceGrad 分支, 现在抛的是 NotImplementedError 而不是 UFLException (ufl.log 模块已经没有了). 真遇到时, 按 Currently no support for ReferenceGrad in CoefficientDerivative 这句原文去搜索仍然有效.

下面的单元保留下来当回归检查用: 哪天它打印出 TEST ERROR, 就说明这条限制又回来了.

Note

Measured on this machine (Firedrake 2026.4.1, UFL 2025.3.0): all four combinations now pass, the cell prints four lines of TEST OK!, and this error no longer shows up on this path.

The restriction itself is still present in UFL, only the exception type has changed: it is raised in the ReferenceGrad branch of ufl/algorithms/apply_derivatives.py, and what is raised now is a NotImplementedError rather than a UFLException (the ufl.log module is gone). Should one run into it, searching for the original wording Currently no support for ReferenceGrad in CoefficientDerivative still works.

The cell below is kept as a regression check: the day it prints TEST ERROR, the restriction is back.

from firedrake import *

def get_s12(mesh, u, v):
    n = FacetNormal(mesh)
    s1 = dot(n, dot(n, grad(grad(grad(u)))))
    s2 = dot(n, dot(n, grad(grad(grad(v)))))
    return s1, s2

def get_s12_v2(mesh, u, v):
    n = FacetNormal(mesh)
    s1 = dot(n, grad(dot(n, grad(grad(u)))))
    s2 = dot(n, grad(dot(n, grad(grad(v)))))
    return s1, s2

def test_grad_n(high_order_mesh=True, fun=get_s12):
    p = 2
    N = 4

    mesh = RectangleMesh(N, N, 1, 1)
        
    if high_order_mesh:
        V = VectorFunctionSpace(mesh, 'CG', 2)
        coords = Function(V).interpolate(mesh.coordinates)
        mesh = Mesh(coords)

    V = FunctionSpace(mesh, 'CG', p)

    u, v = TrialFunction(V), TestFunction(V)

    s1, s2 = fun(mesh, u, v)
    a = inner(grad(u), grad(v))*dx + inner(s1('+'), s2('+'))*dS + inner(s1('-'), s2('-'))*dS
    L = inner(Constant(0), v)*dx

    sol = Function(V)
    prob = LinearVariationalProblem(a, L, sol)
    solver = LinearVariationalSolver(prob)

    solver.solve()

for f in [get_s12, get_s12_v2]:
    for hom in [False, True]:
        try:
            test_grad_n(high_order_mesh=hom, fun=f)
            print(f.__name__, f': High order mesh: {hom}, ', "TEST OK!")
        except Exception as e:
            print(f.__name__, f': High order mesh: {hom}, ', "TEST ERROR: ", e)
get_s12 : High order mesh: False,  TEST OK!
get_s12 : High order mesh: True,  TEST OK!
get_s12_v2 : High order mesh: False,  TEST OK!
get_s12_v2 : High order mesh: True,  TEST OK!

2.1.3. PyErr_Occurred#

python: src/petsc4py.PETSc.c:348918: __Pyx_PyCFunction_FastCall: Assertion `!PyErr_Occurred()' failed.

这类断言失败通常发生在 PETSc 回调 Python 代码的时候: 你写的回调 (自定义预条件子、KSP/SNES 的 monitor、PCPython 等) 里抛了异常, 异常没有被正确地传回 C 端, 于是下一次进入 petsc4py 时断言 !PyErr_Occurred() 就失败了. 报错位置在 petsc4py 生成的 C 文件里, 和你的代码没有直接对应关系, 看那个行号没有用.

排查办法: 把回调函数单独拿出来在 Python 里调用一次, 先保证它自己不出错 —— 最常见的就是变量名拼错、变量没定义这类低级错误. 再配合下一节让 PETSc 打印完整错误信息, 通常能定位到更靠前的出错点.

Assertion failures of this kind usually happen when PETSc calls back into Python code: the callback you wrote (a custom preconditioner, a monitor of KSP/SNES, PCPython and so on) raised an exception, the exception was not propagated back to the C side properly, and so the assertion !PyErr_Occurred() fails the next time petsc4py is entered. The reported location is inside the C file generated by petsc4py and has no direct correspondence to your code, so the line number there is of no use.

How to track it down: take the callback out and call it once from Python on its own, making sure it does not fail by itself — the most common causes are trivial ones such as a misspelled or undefined variable name. Combined with the next section, which makes PETSc print the full error message, this usually pins down an earlier point of failure.

2.1.4. 让 PETSc 打印完整的错误信息 Making PETSc print the full error message#

petsc4py 默认会把 PETSc 的错误压缩成一个 Python 异常, 只保留几行调用栈. 在程序最开始加上

By default petsc4py compresses a PETSc error into a Python exception and keeps only a few lines of the call stack. Adding

from firedrake.petsc import PETSc
PETSc.Sys.popErrorHandler()

就会退回 PETSc 自带的错误处理器, 出错时把完整的 [0]PETSC ERROR: 块打印到标准错误, 里面有 PETSc 版本、configure 选项和 C 端逐层调用栈. 向 Firedrake 或 PETSc 提交 issue 时, 对方需要的正是这一整块信息.

本机实测同一段出错代码, 加与不加的区别:

at the very beginning of the program falls back to PETSc’s own error handler, which prints the complete [0]PETSC ERROR: block to standard error when something goes wrong; it contains the PETSc version, the configure options and the call stack of the C side level by level. When filing an issue against Firedrake or PETSc, this whole block is exactly what is needed.

Measured on this machine with the same failing code, the difference between leaving it out and putting it in (the first block is without it — the petsc4py default; the second is with PETSc.Sys.popErrorHandler()):

# 不加 (petsc4py 默认): 只有压缩过的几行
petsc4py.PETSc.Error: error code 56
[0] VecView() at .../src/vec/vec/interface/vector.c:894
[0] No support for this operation for this object type
[0] No method view for Vec of type (null)
# 加了 PETSc.Sys.popErrorHandler(): PETSc 自己的完整报告
[0]PETSC ERROR: --------------------- Error Message ------------------------------
[0]PETSC ERROR: No support for this operation for this object type
[0]PETSC ERROR: No method view for Vec of type (null)
[0]PETSC ERROR: See https://petsc.org/release/faq/ for trouble shooting.
[0]PETSC ERROR: PETSc Release Version 3.25.0, unknown
[0]PETSC ERROR: <脚本名> with 1 MPI process(es) and PETSC_ARCH ... on <主机名> by <用户名> <时间>
[0]PETSC ERROR: Configure options: ...
[0]PETSC ERROR: #1 VecView() at .../src/vec/vec/interface/vector.c:894

2.1.5. PetscErrorCode#

petsc4py 抛出的 PETSc.Error 带一个 ierr 属性, 就是 PETSc 的错误码; 前面出现的 error code 91、error code 76、error code 56 都是它. 拿到错误码之后对照下表就知道属于哪一类, 常用的几个:

错误码

名字

含义

56

PETSC_ERR_SUP

该对象类型不支持这个操作 (例如对还没设类型的 Vec 调 view)

63

PETSC_ERR_ARG_OUTOFRANGE

参数越界

71

PETSC_ERR_MAT_LU_ZRPVT

LU 分解遇到零主元

76

PETSC_ERR_LIB

PETSc 调用的外部库 (MUMPS、hypre 等) 出错

91

PETSC_ERR_NOT_CONVERGED

求解器没有收敛

在代码里可以按 ierr 分情况处理, 完整例子见 PETSc 一章的 KSP 一节.

手册页: https://petsc.org/release/manualpages/Sys/PetscErrorCode/

The PETSc.Error raised by petsc4py carries an ierr attribute, which is PETSc’s error code; the error code 91, error code 76 and error code 56 seen above are all of this kind. Once the code is known, the table below says which class it belongs to; the common ones are:

Error code

Name

Meaning

56

PETSC_ERR_SUP

the operation is not supported for this object type (for example calling view on a Vec whose type has not been set yet)

63

PETSC_ERR_ARG_OUTOFRANGE

an argument is out of range

71

PETSC_ERR_MAT_LU_ZRPVT

a zero pivot was hit during the LU factorization

76

PETSC_ERR_LIB

an external library called by PETSc (MUMPS, hypre and so on) failed

91

PETSC_ERR_NOT_CONVERGED

the solver did not converge

The code can branch on ierr; for a complete example see the section KSP in the PETSc chapter.

Manual page: https://petsc.org/release/manualpages/Sys/PetscErrorCode/

2.2. 调试 Python 代码 Debugging Python code#

程序抛异常时, 先把回溯的最后几行读完: 找到出错的那一行, 再检查相关变量是不是有异常值 (None、形状不对、nan).

在 Jupyter notebook 里, 出错之后直接执行

When the program raises an exception, first read the last few lines of the traceback to the end: find the line that failed, then check whether the variables involved have unexpected values (None, a wrong shape, nan).

In a Jupyter notebook, running

%debug

就会在出错的那一帧打开 pdb 调试器: p 变量名 打印变量, u / d 在调用栈上下移动, q 退出. 如果希望以后一出错就自动进调试器, 执行一次 %pdb on 即可.

写成脚本运行时对应的做法是: 在可疑位置插一行 breakpoint(), 或者用

right after the error opens the pdb debugger in the frame that failed: p <variable> prints a variable, u / d move up and down the call stack, q quits. To enter the debugger automatically on every future error, run %pdb on once.

The counterpart when running as a script is to insert a breakpoint() line at the suspicious place, or to use

python -m pdb -c continue test.py

让脚本一直跑到出错处再停下.

Firedrake 里有一类异常要特别注意: 报错是 petsc4py.PETSc.Error, Python 回溯只到调用 PETSc 的那一行为止, 真正的原因在 PETSc 的 C 端. 这时 Python 调试器帮不上忙, 应当先按上一节的办法让 PETSc 打印完整错误信息, 不够再上 gdb.

so that the script runs until it fails and stops there.

One class of exceptions in Firedrake deserves particular attention: the error is a petsc4py.PETSc.Error, the Python traceback stops at the line that called PETSc, and the real cause is on PETSc’s C side. A Python debugger is of no help here; one should first make PETSc print the full error message as described in the previous section, and only reach for gdb if that is not enough.

2.3. 调试 C 代码 Debugging C code (gdb)#

Firedrake 的网格管理和线性方程组求解都是 PETSc 做的, 出错也可能出在 PETSc 的 C 代码里. 需要 gdb 这类工具的判据是: 进程死掉或者卡住, 但没有留下 Python 回溯. 典型的有

  • 段错误 (SIGSEGV)、总线错误 (SIGBUS) 之类的信号, 进程直接被杀掉;

  • 程序停在那里不动 (并行时最常见, 见 并行常见陷阱 Common parallel pitfalls), 需要看清它卡在哪个函数上;

  • C 端的断言失败, 例如前面的 PyErr_Occurred.

反过来, 如果 PETSc 是通过错误码返回的, petsc4py 会把它转成 Python 异常, 这时用不上 gdb. 例如下面这个脚本:

Mesh management and the solution of linear systems in Firedrake are both done by PETSc, so a failure may also occur inside PETSc’s C code. A tool such as gdb is needed when the process dies or hangs without leaving a Python traceback. Typical cases are

  • signals such as a segmentation fault (SIGSEGV) or a bus error (SIGBUS), which kill the process outright;

  • the program sitting there doing nothing (most common in parallel, see 并行常见陷阱 Common parallel pitfalls), where one needs to see which function it is stuck in;

  • an assertion failure on the C side, such as the PyErr_Occurred above.

Conversely, when PETSc returns through an error code, petsc4py turns it into a Python exception and gdb is not needed. Take the following script:

# filename: test.py
import sys
import petsc4py
petsc4py.init(sys.argv)
from petsc4py import PETSc
if PETSc.COMM_WORLD.rank == 0:
    PETSc.Vec().create(comm=PETSc.COMM_SELF).view()

它对一个还没设类型的 Vec 调用 view(), 出错信息如下:

It calls view() on a Vec whose type has not been set yet, and the error message is:

$ python test.py
Vec Object: 1 MPI process
  type not yet set
Traceback (most recent call last):
  File "test.py", line 7, in <module>
    PETSc.Vec().create(comm=PETSc.COMM_SELF).view()
  File "PETSc/Vec.pyx", line 140, in petsc4py.PETSc.Vec.view
petsc4py.PETSc.Error: error code 56
[0] VecView() at /home/user/software/firedrake-mini-petsc/src/petsc/src/vec/vec/interface/vector.c:715
[0] No support for this operation for this object type
[0] No method view for Vec of type (null)

Python 回溯是完整的, 一眼就能看出问题出在哪一行. 本章后面的 gdb 命令仍以这个 test.py 作为被调试的程序, 只是为了有个可以照着敲的例子; 真正需要 gdb 的是上面列出的那几种”没有回溯”的场景, 本章 tmux-mpi 的调试命令一节给了一次真实的 SIGBUS 会话记录.

The Python traceback is complete and one can see at a glance which line went wrong. The gdb commands later in this chapter still use this test.py as the program being debugged, simply so that there is an example to type along with; what really calls for gdb are the “no traceback” situations listed above, and the section on the debugging commands of tmux-mpi in this chapter gives the record of a real SIGBUS session.

2.3.1. gdb 命令行说明 The gdb command line#

gdb [options] --args executable-file [inferior-arguments ...]

调试 Firedrake 程序时, 被调试的可执行文件是 Python 解释器本身, 脚本只是它的参数, 所以命令里要写成 --args $(which python) test.py.

When debugging a Firedrake program the executable being debugged is the Python interpreter itself and the script is merely one of its arguments, so the command has to be written as --args $(which python) test.py.

2.3.2. 参数 Arguments (options)#

  1. -x file: 从文件中读取 gdb 命令

  2. -ex COMMAND: 执行一条 gdb 命令

  3. --args exe [exe-args]: 传递参数给 exe

  4. --pid <pid> (或 -p <pid>): 挂到一个正在运行的进程上. 程序挂住不动时用这个, 不必重启程序

  1. -x file: read gdb commands from a file

  2. -ex COMMAND: execute one gdb command

  3. --args exe [exe-args]: pass arguments to exe

  4. --pid <pid> (or -p <pid>): attach to a running process. Use this when the program hangs, so that it does not have to be restarted

2.3.3. gdb 命令 gdb commands#

  1. bt: 查看函数调用栈

  2. run: 运行可执行文件

  3. c: 继续运行

  4. l: 查看代码

  5. p: 打印变量

  6. thread apply all bt: 打印所有线程的调用栈, 排查挂起时常用

  1. bt: show the call stack

  2. run: run the executable

  3. c: continue running

  4. l: list the code

  5. p: print a variable

  6. thread apply all bt: print the call stacks of all threads, commonly used when tracking down a hang

2.3.4. 示例 (调试 test.py) Example (debugging test.py)#

$ gdb  -ex run --args $(which python3) test.py

Note

这套命令在 Linux 和基于 Intel 芯片的 macOS 上可用. Apple Silicon (arm64) 的 macOS 上 gdb 用不了: 本机实测 Homebrew 的 gdb 16.3 连 Mach-O arm64 可执行文件都识别不出来 (报 not in executable format: file format not recognized), set architecture 的候选列表里也没有 aarch64. 这类机器上请改用系统自带的 lldb, 对应命令见后面 tmux-mpi 的调试命令一节.

Note

These commands work on Linux and on Intel-based macOS. On Apple Silicon (arm64) macOS gdb does not work: measured on this machine, Homebrew’s gdb 16.3 does not even recognize a Mach-O arm64 executable (it reports not in executable format: file format not recognized), and aarch64 is not among the candidates of set architecture either. On such machines use the system’s own lldb instead; the corresponding commands are given in the section on the debugging commands of tmux-mpi below.

2.3.5. gdb 控制命令 gdb control commands#

以下命令可以把 gdb 的调试过程保存到文件, 便于贴进 issue.

  1. 输出执行的 gdb 命令 set trace-commands on ref https://sourceware.org/gdb/onlinedocs/gdb/Messages_002fWarnings.html

  2. 关闭分页 set pagination off ref https://sourceware.org/gdb/onlinedocs/gdb/Screen-Size.html

  3. 设置日志文件, 并开启日志 set logging file my.logs, set logging enable on ref https://sourceware.org/gdb/download/onlinedocs/gdb/Logging-Output.html

关掉分页是必要的一步: 默认情况下 gdb 每输出一屏就停下来等按键, 写成脚本自动运行时会一直卡在那里.

The commands below save a gdb session to a file, which makes it easy to paste into an issue.

  1. Echo the gdb commands that are executed: set trace-commands on ref https://sourceware.org/gdb/onlinedocs/gdb/Messages_002fWarnings.html

  2. Turn paging off: set pagination off ref https://sourceware.org/gdb/onlinedocs/gdb/Screen-Size.html

  3. Set a log file and turn logging on: set logging file my.logs, set logging enable on ref https://sourceware.org/gdb/download/onlinedocs/gdb/Logging-Output.html

Turning paging off is a necessary step: by default gdb stops after every screenful of output and waits for a key press, which would make an automated script hang there forever.

2.3.6. gdb 的 python 插件 The python plugin of gdb#

Python 解释器的调用栈在 gdb 里是一堆 _PyEval_EvalFrame 之类的 C 帧, 看不出对应哪一行 Python 代码. gdb 的 python 插件提供了 py-bt 命令, 可以把它翻译成 Python 层的调用栈.

在 ubuntu 上安装 python3-dbg 后, 文件夹 /usr/share/gdb/auto-load/usr/bin/ 下会有如下插件

In gdb the call stack of the Python interpreter is a pile of C frames such as _PyEval_EvalFrame, from which one cannot tell which line of Python code they correspond to. The python plugin of gdb provides the command py-bt, which translates it into a call stack at the Python level.

On ubuntu, after installing python3-dbg, the following plugins are found in the directory /usr/share/gdb/auto-load/usr/bin/

$ ls /usr/share/gdb/auto-load/usr/bin/
python3.10-dbg-gdb.py  python3.10dm-gdb.py  python3.10-gdb.py  python3.10m-gdb.py

其中定义了用于显示 python 调用栈的命令: py-bt. 下面的文件名以 Ubuntu 22.04 的 Python 3.10 为例, 实际使用时按系统里的版本号替换.

gdb 没有自动加载该插件时, 可以手动加载:

They define the command that shows the python call stack: py-bt. The file names below use Python 3.10 of Ubuntu 22.04 as an example; replace them with the version number found on your own system.

When gdb does not load the plugin automatically, it can be loaded by hand:

source /usr/share/gdb/auto-load/usr/bin/python3.10m-gdb.py

或者把这一行写进 gdb 的初始化文件 $HOME/.config/gdb/gdbinit 或当前目录下的 .gdbinit, 以后每次启动都会自动加载.

Alternatively, put this line into gdb’s initialization file $HOME/.config/gdb/gdbinit, or into a .gdbinit in the current directory, and it will be loaded automatically on every start.

2.3.7. 从文件读取 gdb 命令 Reading gdb commands from a file#

命令多了以后一条条敲很麻烦, 更实用的做法是把它们写进一个文件, 让 gdb 自动执行完再退出 —— 挂到一个卡住的进程上抓一份调用栈时尤其方便.

新建文件 cmd.txt 内容如下

Once there are many commands, typing them one by one is tedious; it is more practical to put them into a file and let gdb run them and then exit — this is particularly convenient for attaching to a hung process and grabbing a call stack.

Create a file cmd.txt with the following content

source /usr/share/gdb/auto-load/usr/bin/python3.10m-gdb.py
set trace-commands on
set pagination off
set logging file my.logs
set logging enable on
py-bt
exit

使用 -x 参数运行 gdb, 它会自动运行上面的命令并退出, 结果留在 my.logs 里.

Running gdb with the -x option makes it run the commands above and exit, leaving the result in my.logs.

gdb -x cmd.txt -p <pid>

2.4. 并行程序调试 Debugging parallel programs#

并行下的问题大致分两类, 处理方式完全不同:

  • 某个进程崩溃或报错. 各进程的输出会搅在一起, 而且往往只有一个进程真的出错. 先设法弄清是哪个 rank 出的问题, 再只对那个 rank 开调试器.

  • 程序挂住不动. 最常见的原因是集合操作只被一部分进程调用了 —— 例如把 assemble、solve、norm 或 dat.data 写在 if rank == 0: 里面, 其余进程走不到这一句, 于是所有进程一起等下去. 本项目用 mpiexec -n 2 实测过, 程序会永久挂起, 既不报错也不退出, 详见 并行常见陷阱 Common parallel pitfalls. 这类问题没有任何报错信息, 只能挂上调试器看各个进程分别停在哪里.

不管哪一类, 关键都是给每个进程一个独立的终端, 否则多个调试器的输入输出会互相干扰.

Problems in parallel fall roughly into two classes, which are handled in completely different ways:

  • One process crashes or reports an error. The output of the processes is interleaved, and often only one process really failed. First find out which rank the problem comes from, then open a debugger for that rank only.

  • The program hangs. The most common cause is a collective operation being called by only some of the processes — for example writing assemble, solve, norm or dat.data inside if rank == 0:, so that the other processes never reach this statement and all of them wait forever. This has been measured in this project with mpiexec -n 2: the program hangs permanently, neither erroring out nor exiting, see 并行常见陷阱 Common parallel pitfalls. Problems of this kind come with no error message at all, and the only way is to attach a debugger and see where each process is stopped.

In either case the key is to give every process a terminal of its own, otherwise the input and output of the several debuggers interfere with each other.

2.4.1. PETSc 的参数 The PETSc option -start_in_debugger#

PETSc 自带一个开关: 加上 -start_in_debugger, 它会在初始化阶段给每个进程各开一个终端窗口, 窗口里是已经挂好的调试器.

PETSc comes with a switch of its own: with -start_in_debugger it opens, during initialization, one terminal window per process, each holding a debugger already attached.

mpiexec -n 3 $(which python) test.py  -start_in_debugger

选项要写在脚本名之后: 这样它们才会留在 sys.argv 里, 由 PETSc 在初始化时读走 (Firedrake 一被 import 就会做这件事).

默认行为和平台有关 (以下为 PETSc 3.25 源码中的默认值):

  • 调试器在 PETSc 配置阶段就定下来了: macOS 上默认是 lldb, 其他系统默认是 gdb;

  • 终端在 macOS 上是 Terminal.app, 在其他系统上是 xterm.

几个配套选项, 进程一多就很有用:

选项

作用

-debugger_ranks m,n

只给指定的 rank 开调试器, 默认是所有进程

-start_in_debugger noxterm

不另开窗口, 就在当前终端里挂调试器

-stop_for_debugger

不开调试器, 只打印”如何手工挂上去”的提示, 然后停下来等

-debug_terminal <cmd>

换一个终端程序, 例如 -debug_terminal "screen -X -S debug screen"

-debugger_pause <秒>

等若干秒再挂上去, 适合终端启动慢的场合

其中 -stop_for_debugger 在只有 ssh、没有 X 转发的机器上尤其有用: 开不出窗口时, 让它打印出进程号, 再自己 gdb -p <pid> 挂上去.

参考:

  1. https://petsc.org/main/manualpages/Sys/PetscInitialize/

  2. https://petsc.org/main/manualpages/Sys/PetscAttachDebugger/

  3. https://petsc.org/main/manualpages/Sys/PetscSetDebugTerminal/

The options have to be written after the script name: only then do they stay in sys.argv and get picked up by PETSc during initialization (which happens as soon as Firedrake is imported).

The default behavior depends on the platform (the defaults below are those in the PETSc 3.25 sources):

  • the debugger is fixed at PETSc configure time: on macOS the default is lldb, on other systems it is gdb;

  • the terminal is Terminal.app on macOS and xterm elsewhere.

A few companion options, which become useful as soon as there are many processes:

Option

Effect

-debugger_ranks m,n

open a debugger only for the given ranks; the default is all processes

-start_in_debugger noxterm

do not open a separate window, attach the debugger in the current terminal

-stop_for_debugger

do not open a debugger, only print a hint on how to attach one by hand, then stop and wait

-debug_terminal <cmd>

use a different terminal program, for example -debug_terminal "screen -X -S debug screen"

-debugger_pause <seconds>

wait a number of seconds before attaching, suitable when the terminal starts slowly

Among these, -stop_for_debugger is particularly useful on a machine reachable only over ssh with no X forwarding: when no window can be opened, let it print the process id and then attach with gdb -p <pid> oneself.

References:

  1. https://petsc.org/main/manualpages/Sys/PetscInitialize/

  2. https://petsc.org/main/manualpages/Sys/PetscAttachDebugger/

  3. https://petsc.org/main/manualpages/Sys/PetscSetDebugTerminal/

Tip

修改 xterm 窗口显示效果 (终端是 xterm 时才有意义, 即 Linux)

参考: http://www.futurile.net/2016/06/14/xterm-setup-and-truetype-font-configuration/

$ cat ~/.Xdefaults
xterm*faceName: Monospace
xterm*faceSize: 12
xterm*foreground: rgb:a8/a8/a8
xterm*background: rgb:00/00/00

Tip

Changing the appearance of the xterm window (only meaningful when the terminal is xterm, so on Linux)

Reference: http://www.futurile.net/2016/06/14/xterm-setup-and-truetype-font-configuration/

$ cat ~/.Xdefaults
xterm*faceName: Monospace
xterm*faceSize: 12
xterm*foreground: rgb:a8/a8/a8
xterm*background: rgb:00/00/00

2.4.2. 工具 The tool tmux-mpi#

-start_in_debugger 要能开出图形窗口, 在只有 ssh 终端的集群上通常用不了. tmux-mpi 换了个思路: 它给每个 MPI 进程开一个 tmux 窗口, 每个窗口里跑一个调试器, 全程在字符终端里完成, 不需要 X 转发, 断线之后重新 attach 还能接着调.

参考:

  1. Firedrake wiki 上的说明: firedrakeproject/firedrake

  2. Tips of Firedrake Wiki: firedrakeproject/firedrake

  3. tmux-mpi 配置与高级用法: wrs20/tmux-mpi

-start_in_debugger needs to be able to open graphical windows, which usually rules it out on a cluster with nothing but an ssh terminal. tmux-mpi takes a different route: it opens one tmux window per MPI process, each running a debugger, entirely inside a character terminal, with no X forwarding required, and after a dropped connection one can attach again and carry on debugging.

References:

  1. The description on the Firedrake wiki: firedrakeproject/firedrake

  2. Tips of Firedrake Wiki: firedrakeproject/firedrake

  3. Configuration and advanced usage of tmux-mpi: wrs20/tmux-mpi

2.4.2.1. 安装 tmux-mpi Installing tmux-mpi#

tmux-mpi 依赖 tmux 和 dtach 两个外部程序, 需要先装好.

tmux-mpi depends on the two external programs tmux and dtach, which have to be installed first.

  1. 安装 tmuxInstalling tmux

    Ubuntu:

    sudo apt-get install tmux
    

    macOS:

    brew install tmux
    
  2. 安装 dtach (tmux-mpi 依赖)Installing dtach (a dependency of tmux-mpi)

    dtach 一般不在发行版仓库里, 需要自己编译, 然后把二进制文件拷到某个在 PATH 中的目录, 如 $HOME/bin.dtach is usually not in the distribution repositories and has to be compiled by hand; the binary is then copied into a directory on the PATH, such as $HOME/bin.

    git clone https://github.com/crigler/dtach
    cd dtach
    ./configure
    make
    mkdir -p $HOME/bin
    cp dtach $HOME/bin
    export PATH=$PATH:$HOME/bin
    

    运行 which dtach 确认安装是否成功. 最后那行 export 只对当前终端有效, 要长期生效需写进 ~/.bashrc 或 ~/.zshrc.Run which dtach to confirm that the installation succeeded. The last export line only affects the current terminal; to make it permanent write it into ~/.bashrc or ~/.zshrc.

  3. 安装 tmux-mpiInstalling tmux-mpi

    在 Firedrake 的虚拟环境里用 pip 安装Install it with pip inside Firedrake’s virtual environment

    pip install --upgrade --no-cache-dir git+https://github.com/wrs20/tmux-mpi@master
    

    也可以用 pipx install git+https://github.com/wrs20/tmux-mpi@master 装到全局, 不污染虚拟环境. 装好后 which tmux-mpi 应当能找到它.It can also be installed globally with pipx install git+https://github.com/wrs20/tmux-mpi@master, which leaves the virtual environment untouched. After installation which tmux-mpi should find it.

2.4.2.2. 调试命令 Debugging commands#

Linux 和基于 Intel 芯片的 macOS 系统可以使用 gdb 作为调试器; 在使用 Apple Silicon (arm64) 的 macOS 系统上, 需改用 lldb.

On Linux and on Intel-based macOS gdb can be used as the debugger; on macOS with Apple Silicon (arm64) one has to use lldb instead.

  1. 启动调试器Start the debugger

    tmux-mpi 3 gdb -ex run --args $(which python) test.py
    

    用 lldb 时对应写成With lldb the corresponding form is

    tmux-mpi 3 lldb -o run -o continue -- $(which python) test.py
    

    这里比 gdb 多一条 -o continue: 本机实测 lldb -o run 会先停在 dyld_start (动态加载器的入口), 要再 continue 一次程序才真正跑起来.There is one more -o continue here than with gdb: measured on this machine, lldb -o run first stops at dyld_start (the entry point of the dynamic loader), and only after one more continue does the program really start running.

  2. Attach 到相应的伪终端, 每个进程一个窗口. (这里是 tmux 的一个 session, 有多个 window)Attach to the corresponding pseudo-terminals, one window per process. (This is a single tmux session with several windows.)

    tmux attach -t tmux-mpi
    

    tmux 的窗口切换: Ctrl-b 之后按 n / p 切到下一个 / 上一个窗口, 按数字键直接跳到对应编号的进程, 按 d 从 session 里退出 (进程还在跑, 之后可以再 attach 回来).Switching windows in tmux: after Ctrl-b, press n / p to move to the next / previous window, press a digit to jump straight to the process with that number, and press d to detach from the session (the processes keep running and can be attached to again later).

  3. 使用 gdb 调试命令调试, 例如查看栈回溯 (bt)、反汇编 (disassemble)、寄存器和内存映射. 下面是一次实际调试 SIGBUS 错误的会话记录: 程序在 MPICH 的 memcpy 里收到总线错误, 于是先用 bt 看是从哪里调用过来的, 再打印相关寄存器和 info proc mapping, 判断拷贝的目标地址是不是落到了共享内存段 (/dev/shm/mpich_shar_tmp...) 的边界之外.Debug with the gdb commands, for example looking at the backtrace (bt), the disassembly (disassemble), the registers and the memory mappings. Below is the record of a real session debugging a SIGBUS error: the program received a bus error inside MPICH’s memcpy, so bt was used first to see where the call came from, and then the relevant registers and info proc mapping were printed in order to decide whether the destination address of the copy had fallen outside the bounds of the shared memory segment (/dev/shm/mpich_shar_tmp...).

    Thread 1 "python" received signal SIGBUS, Bus error.
    __memmove_evex_unaligned_erms () at ../sysdeps/x86_64/multiarch/memmove-vec-unaligned-erms.S:708
    708     ../sysdeps/x86_64/multiarch/memmove-vec-unaligned-erms.S: No such file or directory.
    (gdb) bt
    #0  __memmove_evex_unaligned_erms ()
        at ../sysdeps/x86_64/multiarch/memmove-vec-unaligned-erms.S:708
    #1  0x00007feb09870bfb in memcpy (__len=47424, __src=<optimized out>,
        __dest=<optimized out>)
        at /usr/include/x86_64-linux-gnu/bits/string_fortified.h:29
    #2  MPID_nem_mpich_sendv_header (again=<synthetic pointer>, vc=0x55a9d0d49bb0,
        n_iov=<synthetic pointer>, iov=<synthetic pointer>)
        at ./src/mpid/ch3/channels/nemesis/include/mpid_nem_inline.h:396
    #3  MPIDI_CH3_iSendv (vc=0x55a9d0d49bb0, sreq=<optimized out>,
        sreq@entry=0x55a9d3b5e430, iov=iov@entry=0x7ffe5ee5e0f0,
        n_iov=n_iov@entry=2) at src/mpid/ch3/channels/nemesis/src/ch3_isendv.c:69
    #4  0x00007feb0985c7dc in MPIDI_CH3_EagerContigIsend (
        sreq_p=sreq_p@entry=0x7ffe5ee5e268,
        reqtype=reqtype@entry=MPIDI_CH3_PKT_EAGER_SEND, buf=<optimized out>,
        data_sz=<optimized out>, rank=rank@entry=2, tag=<optimized out>,
        comm=0x55a9d327aab0, context_offset=0) at src/mpid/ch3/src/ch3u_eager.c:545
        ...
    (gdb) disassemble
    Dump of assembler code for function __memmove_evex_unaligned_erms:
    0x00007feb0c349f40 <+0>:     endbr64
    0x00007feb0c349f44 <+4>:     mov    %rdi,%rax
    0x00007feb0c349f47 <+7>:     cmp    $0x20,%rdx
    0x00007feb0c349f4b <+11>:    jb     0x7feb0c349f80 <__memmove_evex_unaligned_erms+64>
    0x00007feb0c349f4d <+13>:    vmovdqu64 (%rsi),%ymm16
    0x00007feb0c349f53 <+19>:    cmp    $0x40,%rdx
    0x00007feb0c349f57 <+23>:    ja     0x7feb0c34a000 <__memmove_evex_unaligned_erms+192>
        ...
    0x00007feb0c34a286 <+838>:   add    %rdi,%rsi
    0x00007feb0c34a289 <+841>:   sub    %rdi,%rcx
    => 0x00007feb0c34a28c <+844>:   rep movsb %ds:(%rsi),%es:(%rdi)
    0x00007feb0c34a28e <+846>:   vmovdqu64 %ymm16,(%r8)
    0x00007feb0c34a294 <+852>:   vmovdqu64 %ymm17,0x20(%r8)
        ...
    --Type <RET> for more, q to quit, c to continue without paging--q
    Quit
    (gdb) p/x $rsi
    $13 = 0x7fe9f43c93a0
    (gdb) p/x $rdi
    $14 = 0x7feb016a4000
    (gdb) p/x $rcx
    $15 = 0x25b0
    (gdb) p $rcx
    $16 = 9648
    (gdb) p/x $rdi + $rcx
    $17 = 0x7feb016a65b0
    (gdb) p/x $rsi
    $13 = 0x7fe9f43c93a0
    (gdb) p/x $rdi
    $14 = 0x7feb016a4000
    (gdb) info proc mapping
        ...
        0x7feafe633000     0x7feafe634000     0x1000     0x6000  r--p   /usr/lib/python3.10/lib-dynload/_bz2.cpython-310-x86_64-linux-gnu.so
        0x7feafe634000     0x7feafe635000     0x1000     0x7000  rw-p   /usr/lib/python3.10/lib-dynload/_bz2.cpython-310-x86_64-linux-gnu.so
        0x7feafe635000     0x7feafebf7000   0x5c2000        0x0  rw-p
        0x7feafebf7000     0x7feb03afc000  0x4f05000        0x0  rw-s   /dev/shm/mpich_shar_tmpMCilHL (deleted)
        0x7feb03afc000     0x7feb03bfc000   0x100000        0x0  rw-p
        ...
    

    lldb 里对应的命令是 bt (相同)、thread backtrace all、disassemble、register read 和 image list.The corresponding commands in lldb are bt (the same), thread backtrace all, disassemble, register read and image list.

2.5. 编译 .pyx 文件 Compiling .pyx files#

2.5.1. Firedrake 中的 .pyx 文件 The .pyx files in Firedrake#

Firedrake 有一部分底层代码是用 Cython 写的, 放在 firedrake/cython/ 目录下, 编译成扩展模块 (.so) 之后才能被 import:

Part of Firedrake’s low-level code is written in Cython and lives in the directory firedrake/cython/; it can only be imported once it has been compiled into extension modules (.so):

dmcommon.pyx             DM (DMPlex) 相关的公共工具函数
extrusion_numbering.pyx  拉伸网格的自由度编号
hdf5interface.pyx        HDF5 底层接口
mgimpl.pyx               多重网格的底层编号
patchimpl.pyx            PCPatch 的回调
spatialindex.pyx         空间索引 (点定位)
supermeshimpl.pyx        supermesh 投影

改动这些文件之后必须重新编译, 否则改的是 .pyx, 跑的仍然是旧的 .so. 这件事只对源码 (editable) 安装有意义: 用 wheel 装的 Firedrake 目录里虽然也带着 .pyx, 但那只是随包附带的源文件, 改了不会生效.

在 Firedrake 源码树的根目录 (有 setup.py 的那一层) 执行下面的命令就地重新编译:

The annotations in the listing give, in order: common utilities for DM (DMPlex), the numbering of degrees of freedom on extruded meshes, the low-level HDF5 interface, the low-level numbering for multigrid, the callbacks of PCPatch, the spatial index (point location) and supermesh projection.

After changing these files they must be recompiled, otherwise what was changed is the .pyx while what runs is still the old .so. This only makes sense for a source (editable) installation: a Firedrake installed from a wheel also carries the .pyx files in its directory, but those are merely the source files shipped with the package, and changing them has no effect.

Run the following command in the root of the Firedrake source tree (the level that contains setup.py) to recompile in place:

python setup.py build_ext --inplace

也可以按官方文档的方式重装一次, 效果相同:

Reinstalling once in the way the official documentation describes has the same effect:

pip install --no-build-isolation --editable .

Note

本机是非 editable 安装, 没有 Firedrake 源码树, 上面两条命令未在本机实测. 已经确认的是: Firedrake 仓库根目录仍有 setup.py (setuptools + Cython 构建), 官方安装文档给出的开发安装方式是 pip install --no-build-isolation --editable ...; 上面的 .pyx 文件清单来自 Firedrake 2026.4.1 的安装目录.

Note

This machine has a non-editable installation and no Firedrake source tree, so the two commands above have not been tested on this machine. What has been confirmed is: the root of the Firedrake repository still contains a setup.py (a setuptools + Cython build), the development installation described in the official installation documentation is pip install --no-build-isolation --editable ..., and the list of .pyx files above comes from the installation directory of Firedrake 2026.4.1.

2.6. 排查清单 Checklist#

卡住的时候可以按下面这份通用清单逐条过一遍.

  1. 先把报错读完. Python 回溯看最后几行; PETSc 的错误块看 [0]PETSC ERROR: 的第一条, 后面几条只是调用栈.

  2. 变量是否初始化. 尤其是循环里反复使用的 Function, 每一步之前是否需要清零.

  3. 下标是否错位. n 还是 n+1 还是 n-1, 各自对应哪个时刻的值, 算一遍对上再往下写.

  4. 把规模缩小. 网格调到 4x4、时间步只跑两步, 出错更快, 输出也读得完.

  5. 先串行, 后并行. 串行跑不通的程序不要拿去并行; 并行结果对不对, 用集合量 (assemble、norm、errornorm) 和串行结果逐位对照.

  6. 确认参数真的生效了. 求解器参数拼错不会报错, 只会被忽略 (本机实测把 ksp_type 误写成 ksp_typ, 程序照常跑完). 用 ksp_view 看实际用的是什么求解器, 加命令行参数 -options_left 列出所有没被用上的选项.

When stuck, go through this general checklist item by item.

  1. Read the error message to the end first. For a Python traceback look at the last few lines; for a PETSc error block look at the first [0]PETSC ERROR: line, the following ones are only the call stack.

  2. Are the variables initialized? Especially a Function reused inside a loop: does it have to be zeroed before every step?

  3. Are the indices off by one? Is it n, n+1 or n-1, and which time level does each of them stand for? Work it out once and check it before writing on.

  4. Shrink the problem. Turn the mesh down to 4x4 and run only two time steps: the failure comes sooner and the output can be read in full.

  5. Serial first, parallel afterwards. Do not take a program that fails in serial into parallel; to check whether a parallel result is right, compare collective quantities (assemble, norm, errornorm) digit by digit against the serial result.

  6. Confirm that the options really took effect. A misspelled solver parameter raises no error and is simply ignored (measured on this machine: ksp_type mistyped as ksp_typ and the program ran to completion as usual). Use ksp_view to see which solver is actually used, and add the command-line option -options_left to list all the options that were not used.