Skip to content

Zero-allocation tangent for ViscousPolyconvex - #232

Open
miguelmaso wants to merge 1 commit into
mainfrom
fix-piracy
Open

miguelmaso wants to merge 1 commit into
mainfrom
fix-piracy

Conversation

@miguelmaso

Copy link
Copy Markdown
Collaborator

Add generated implementations of - (binary and unary) for TensorValue, mirroring the existing +. Without them, subtraction of fourth-order tensors fell back to the generic Gridap implementation, which exceeds Julia's inlining threshold for TensorValue{9,9} and heap-allocates the result.

Those subtractions appear in ∂Sv∂C_Cᵥfix and ∂Cv⁻¹∂C, i.e. exactly on the path of the viscoelastic tangent operator, which is why only ∂∂Ψ∂FF allocated while Ψ and ∂Ψ∂F did not.

+ is also generalized from TensorValue{D,D} to TensorValue{D1,D2} so that non-square tensors are covered as well. No change in the mathematics.

The three methods build the result with an explicit element type, which also covers empty tensors: Gridap builds TensorValue{0,3} values when it writes the vertex grid of a 3D model, and leaving the type implicit made them throw UndefVarError: T not defined, breaking writevtk(model) for any code that loads HyperFEM. The price of fixing the element type with promote_type is that adding two Bool tensors now errors instead of promoting to Int as Gridap does; the tensors involved in a finite element evaluation are floating point.

Visco-polyconvex ∂∂Ψ∂FF: 187 allocs / 15552 B / 5367 ns -> 0 allocs / 0 B / 599 ns.
Visco-elastic ∂∂Ψ∂FF: 1021 allocs / 49600 ns -> 501 allocs / 34200 ns.

Also interpolate the arguments of the tensor algebra benchmarks. Without $, the benchmarked expressions read non-const globals, so the call is dynamically dispatched and the result is boxed: this is the source of the spurious "1 alloc of 0.656 kB" reported for IIsym, push_forward_C_to_F and ×ᵢ⁴, which are already allocation-free.

See #231

Add generated implementations of `-` (binary and unary) for `TensorValue`,
mirroring the existing `+`. Without them, subtraction of fourth-order tensors
fell back to the generic Gridap implementation, which exceeds Julia's inlining
threshold for `TensorValue{9,9}` and heap-allocates the result.

Those subtractions appear in `∂Sv∂C_Cᵥfix` and `∂Cv⁻¹∂C`, i.e. exactly on the
path of the viscoelastic tangent operator, which is why only `∂∂Ψ∂FF` allocated
while `Ψ` and `∂Ψ∂F` did not.

`+` is also generalized from `TensorValue{D,D}` to `TensorValue{D1,D2}` so that
non-square tensors are covered as well. No change in the mathematics.

The three methods build the result with an explicit element type, which also
covers empty tensors: Gridap builds `TensorValue{0,3}` values when it writes the
vertex grid of a 3D model, and leaving the type implicit made them throw
`UndefVarError: T not defined`, breaking `writevtk(model)` for any code that
loads HyperFEM. The price of fixing the element type with `promote_type` is that
adding two `Bool` tensors now errors instead of promoting to `Int` as Gridap
does; the tensors involved in a finite element evaluation are floating point.

Visco-polyconvex ∂∂Ψ∂FF: 187 allocs / 15552 B / 5367 ns -> 0 allocs / 0 B / 599 ns.
Visco-elastic ∂∂Ψ∂FF: 1021 allocs / 49600 ns -> 501 allocs / 34200 ns.

Also interpolate the arguments of the tensor algebra benchmarks. Without `$`,
the benchmarked expressions read non-const globals, so the call is dynamically
dispatched and the result is boxed: this is the source of the spurious "1 alloc
of 0.656 kB" reported for `IIsym`, `push_forward_C_to_F` and `×ᵢ⁴`, which are
already allocation-free.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@miguelmaso

Copy link
Copy Markdown
Collaborator Author

#232 is fixing the issues introduced by #231: breaking writevtk due to type pyracy

⚠️⚠️Type piracy must be addressed⚠️⚠️

@github-actions

Copy link
Copy Markdown

Benchmark Results (Julia v1)

Time benchmarks
main d25fca3... main / d25fca3...
Constitutive models/Visco-elastic Ψ 0.0379 ± 0.0058 ms 18.4 ± 7.5 μs 2.06 ± 0.89
Constitutive models/Visco-elastic ∂Ψ∂F 0.0445 ± 0.0074 ms 25.2 ± 8.9 μs 1.77 ± 0.69
Constitutive models/Visco-elastic ∂∂Ψ∂FF 0.0785 ± 0.021 ms 0.0514 ± 0.016 ms 1.53 ± 0.63
Constitutive models/Visco-polyconvex Ψ 0.161 ± 0.01 μs 0.16 ± 0.001 μs 1.01 ± 0.063
Constitutive models/Visco-polyconvex ∂Ψ∂F 0.181 ± 0.01 μs 0.19 ± 0.001 μs 0.953 ± 0.053
Constitutive models/Visco-polyconvex ∂∂Ψ∂FF 11.2 ± 5.6 μs 0.781 ± 0.011 μs 14.3 ± 7.2
Simulations/StaticMechanicalDirichlet 0.155 ± 0.0073 s 0.161 ± 0.0097 s 0.965 ± 0.074
Simulations/StaticMechanicalNeumann 0.122 ± 0.013 s 0.126 ± 0.016 s 0.968 ± 0.16
Simulations/ViscoElastic 10.3 s 7.29 s 1.41
Tensor algebra/Cofactor 0.07 ± 0.01 μs 0.08 ± 0 μs 0.875 ± 0.12
Tensor algebra/Det(A)Inv(A') 0.1 ± 0.01 μs 0.12 ± 0.011 μs 0.833 ± 0.11
Tensor algebra/IIsym 0.06 ± 0.01 μs 0.06 ± 0 μs 1 ± 0.17
Tensor algebra/push_forward_C_to_F 0.27 ± 0.019 μs 0.211 ± 0.011 μs 1.28 ± 0.11
Tensor algebra/×ᵢ⁴ 0.051 ± 0.01 μs 0.06 ± 0 μs 0.85 ± 0.17
Tensor algebra/δδ_λ_2d 20 ± 10 ns 20 ± 10 ns 1 ± 0.71
Tensor algebra/δδ_μ_2d 20 ± 10 ns 20 ± 10 ns 1 ± 0.71
time_to_load 1.96 ± 0.0092 s 2.03 ± 0.0093 s 0.966 ± 0.0063
Memory benchmarks
main d25fca3... main / d25fca3...
Constitutive models/Visco-elastic Ψ 0.56 k allocs: 0.044 MB 0.212 k allocs: 26.8 kB 1.68
Constitutive models/Visco-elastic ∂Ψ∂F 0.609 k allocs: 0.0489 MB 0.261 k allocs: 31.7 kB 1.58
Constitutive models/Visco-elastic ∂∂Ψ∂FF 1.03 k allocs: 0.0829 MB 0.508 k allocs: 0.0573 MB 1.45
Constitutive models/Visco-polyconvex Ψ 0 allocs: 0 B 0 allocs: 0 B
Constitutive models/Visco-polyconvex ∂Ψ∂F 0 allocs: 0 B 0 allocs: 0 B
Constitutive models/Visco-polyconvex ∂∂Ψ∂FF 0.187 k allocs: 15.2 kB 0 allocs: 0 B
Simulations/StaticMechanicalDirichlet 1.51 M allocs: 0.108 GB 1.51 M allocs: 0.108 GB 1
Simulations/StaticMechanicalNeumann 1.47 M allocs: 0.0922 GB 1.47 M allocs: 0.0922 GB 1
Simulations/ViscoElastic 0.16 G allocs: 12.5 GB 0.0801 G allocs: 8.52 GB 1.46
Tensor algebra/Cofactor 2 allocs: 0.156 kB 2 allocs: 0.156 kB 1
Tensor algebra/Det(A)Inv(A') 4 allocs: 0.25 kB 4 allocs: 0.25 kB 1
Tensor algebra/IIsym 1 allocs: 0.656 kB 1 allocs: 0.656 kB 1
Tensor algebra/push_forward_C_to_F 1 allocs: 0.656 kB 1 allocs: 0.656 kB 1
Tensor algebra/×ᵢ⁴ 1 allocs: 0.656 kB 1 allocs: 0.656 kB 1
Tensor algebra/δδ_λ_2d 0 allocs: 0 B 0 allocs: 0 B
Tensor algebra/δδ_μ_2d 0 allocs: 0 B 0 allocs: 0 B
time_to_load 0.205 k allocs: 11.9 kB 0.205 k allocs: 11.9 kB 1

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant