From patchwork Wed Jul 8 09:56:41 2026 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Philipp Tomsich X-Patchwork-Id: 138722 Return-Path: X-Original-To: patchwork@sourceware.org Delivered-To: patchwork@sourceware.org Received: from vm01.sourceware.org (localhost [IPv6:::1]) by sourceware.org (Postfix) with ESMTP id 3EB464BA23C3 for ; Wed, 8 Jul 2026 09:57:26 +0000 (GMT) DKIM-Filter: OpenDKIM Filter v2.11.0 sourceware.org 3EB464BA23C3 Authentication-Results: sourceware.org; dkim=pass (2048-bit key, unprotected) header.d=vrull.eu header.i=@vrull.eu header.a=rsa-sha256 header.s=google header.b=HhiPamvx X-Original-To: gcc-patches@gcc.gnu.org Delivered-To: gcc-patches@gcc.gnu.org Received: from mail-lj1-x22d.google.com (mail-lj1-x22d.google.com [IPv6:2a00:1450:4864:20::22d]) by sourceware.org (Postfix) with ESMTPS id 651924BA2E1B for ; Wed, 8 Jul 2026 09:56:47 +0000 (GMT) DMARC-Filter: OpenDMARC Filter v1.4.2 sourceware.org 651924BA2E1B Authentication-Results: sourceware.org; dmarc=none (p=none dis=none) header.from=vrull.eu Authentication-Results: sourceware.org; spf=pass smtp.mailfrom=vrull.eu ARC-Filter: OpenARC Filter v1.0.0 sourceware.org 651924BA2E1B Authentication-Results: sourceware.org; arc=none smtp.remote-ip=2a00:1450:4864:20::22d ARC-Seal: i=1; a=rsa-sha256; d=sourceware.org; s=key; t=1783504607; cv=none; b=roRmihBkugZCfxDVvSplTkLIswMErWI3DmUwoTyGjXFEBV01j1Vn1pHTbQfLLAQS8TDzVhHmgZJcbbSNYpKfY4VYNJROYRfQiBw4K4cpRcR+w7xIviy1jooC3plWYStcjiS6/OA1lxzpGS4CFjlbO0sGB/4xgbhWMAZL2Fp7GGc= ARC-Message-Signature: i=1; a=rsa-sha256; d=sourceware.org; s=key; t=1783504607; c=relaxed/simple; bh=2wKs1YGlBOpu8Q6+bLjV0sothmv46TZH1LyA3UbqnVw=; h=DKIM-Signature:From:To:Subject:Date:Message-Id:MIME-Version; b=Y9wNkgHvu8P7YmcBtou3/gGibKKFGH5VF8nBLxPCpZr0fXGJGBJ6jIJhd4GD25Gbc5rJKtH7DmlMI8KUwx9iAATIA1oleSxEX3JMDRcrAM5n/2HpWeYLDa+squ7lp2Fo+nWj5LtO2H4eVa4FHyiaP1ZsmX08BiwCCzd144p3k8k= ARC-Authentication-Results: i=1; sourceware.org; dkim=pass (2048-bit key, unprotected) header.d=vrull.eu header.i=@vrull.eu header.a=rsa-sha256 header.s=google header.b=HhiPamvx DKIM-Filter: OpenDKIM Filter v2.11.0 sourceware.org 651924BA2E1B Received: by mail-lj1-x22d.google.com with SMTP id 38308e7fff4ca-39c68f1e863so91311fa.3 for ; Wed, 08 Jul 2026 02:56:47 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=vrull.eu; s=google; t=1783504606; x=1784109406; darn=gcc.gnu.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=Bp3/jtaGvUoI0IYOFwJi1/aG+4seCSmdnGsu0fOvbAY=; b=HhiPamvxpGE707MB4Dp9R0HjYKgwY+vvjTATn3h5i7htQn67fU89aO/E38z3Ny6cOa F2PcsdwUXSl1hCQ4hKoSX4RMlmFsvxbPaAC3igVzRDC6QQopDyxzQqfsC8s6P6JeJ74H YPTdSyvBCjXqzYY/piWH25nufZtfexiUIfU+u5doE3ns7C9bgM5iizjR736SqebIVgQz g+0izLzVyUOaR9xMwasoUnubfDjOJApBFHZemxGhAQvyDKVh2kMhBw1uX8Clld0q2C0K gCX+Jh7y1OPOXvMk9AYiqa4Xok8ifdR3A5wtFRvpKO3xqW0AA6+N5WYC7bwmqMRjxQU8 GpaQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783504606; x=1784109406; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Bp3/jtaGvUoI0IYOFwJi1/aG+4seCSmdnGsu0fOvbAY=; b=JruY6097jRAXPRbKNu6UthCMvjXh7dWjWslmF1PMKNOe+vr6ti8qdlDQAEH9KbSuLX srZPZX0K3l9TjxiHWezArsu46eZ4/vxZxAK8J4CHswHEvb/ezzS9/CSHISbSGo08Cu7d nw1SpxeFdEiFp/pdNptfi0vapJRz2ygANuVHGDhU+tNLqCnbjT1hXYGETrbnOeOAX6OB upbPO/Mg+enqjB3qwfu8r/ld+1blSNE6d/BHZm70DvvJ4dP5/WIXxqr5JJ8HnJV+S61u 2endSDUuPLsVlCJ95lPb47NbunCErRkmahvulAumorNnCpg9sIZcWP0Gn3F37Gk8+rsS EbWQ== X-Gm-Message-State: AOJu0YzL1MLTpe21fg5iSjAVSPmlPgyj6qKxZwnna+o7dD29CjVSVQJQ +d6TapMwo7CYCBbpuYQ6ldaktBe0yMljvXE2Lp8vms67Nz6rqQbCtOkVB2BIIB1StJdHSkMATzw bmm2UJxc= X-Gm-Gg: AfdE7cm049GuSx5kCcLAzcYcJcpbkP3X62VTCucVpT+Mwaxk5b/EWldDU9CTUaer/Rm v7JYT8LeVf8rliRV7Y+g9hCxfcdhrjuGqBTUg2cD0Rlkb3BEPwYyJfhSgmqMabTtKL0XG0ENQe+ VnY7mqdZ6VAFEt/E5NZHf2Pq0v6hLBrD4GrXVLEoSrzRWLBBw9nKq7ITA3W8HdXICHh8qY+IgLY TEUOclQHUkU2C8VgZZdxO2E+s+LmvTBojZvmKdpUOjgujJmxDf2a7ChVOoihi5X0AZr4xpuN7yo cysNRYO9GhsfTkxLuOAyThUgsSX8aeBYBk5mBAzmZT4xotvN1hx0DqEF1jwvDSnde/hnnIuw5hd lnJuuTxGIihJYxcZEkKQNRe/HDZ2IgWE+MmjPz1ObYNpUypFSff91tz7aeDXRi/etjGJjpVh/f7 StgVflk8ytXNHU6Q2LhFfqaGKTdFnCVDM= X-Received: by 2002:a2e:b2d4:0:b0:39c:78b8:6056 with SMTP id 38308e7fff4ca-39c79b12569mr1638541fa.7.1783504605846; Wed, 08 Jul 2026 02:56:45 -0700 (PDT) Received: from ubuntu-focal.. ([2a01:4f9:3a:1e26::2]) by smtp.gmail.com with ESMTPSA id 38308e7fff4ca-39b4ad1c8f9sm33614971fa.4.2026.07.08.02.56.45 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 08 Jul 2026 02:56:45 -0700 (PDT) From: Philipp Tomsich To: gcc-patches@gcc.gnu.org Cc: Philipp Tomsich Subject: [PATCH v2] tree-ssa-math-opts: separate FMA-deferring trigger from reassoc reorder Date: Wed, 8 Jul 2026 11:56:41 +0200 Message-Id: <20260708095641.1611646-1-philipp.tomsich@vrull.eu> X-Mailer: git-send-email 2.34.1 MIME-Version: 1.0 X-Spam-Status: No, score=-12.0 required=5.0 tests=BAYES_00, DKIM_SIGNED, DKIM_VALID, DKIM_VALID_AU, DKIM_VALID_EF, GIT_PATCH_0, KAM_SHORT, PROLO_LEO1, RCVD_IN_DNSWL_NONE, SPF_HELO_NONE, SPF_PASS, TXREP shortcircuit=no autolearn=ham autolearn_force=no version=3.4.6 X-Spam-Checker-Version: SpamAssassin 3.4.6 (2021-04-09) on sourceware.org X-BeenThere: gcc-patches@gcc.gnu.org X-Mailman-Version: 2.1.30 Precedence: list List-Id: Gcc-patches mailing list List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: gcc-patches-bounces~patchwork=sourceware.org@gcc.gnu.org param_avoid_fma_max_bits gates two independent transforms: the widening_mul FMA-deferring in tree-ssa-math-opts.cc, which leaves a loop-carried multiply-add as fmul + fadd, and the reassoc loop-carried-FMA reorder added in r14-5779-g746344dd538 (PR tree-optimization/110279), which parallelises 3+ operand chains. Both fire when TYPE_SIZE (elt) <= avoid-fma-max-bits, so for a given type they switch on at the same threshold and a target cannot keep one while dropping the other. This hurts the AArch64 AVOID_CROSS_LOOP_FMA cores (the Ampere-1 family): a 2-operand reduction such as an sgemm inner K-loop is left as fmul + fadd, slower than fmadd on their pipeline, yet setting avoid-fma-max-bits to 0 to avoid it also disables the reorder. Add a new --param=widening-mul-defer-fma (default 1) that gates only the widening_mul deferring; avoid-fma-max-bits keeps gating the reorder alone. The AVOID_CROSS_LOOP_FMA callback sets the param to 0. No other target and no default behaviour changes. gcc/ChangeLog: * params.opt (-param=widening-mul-defer-fma=): New param. * tree-ssa-math-opts.cc (math_opts_dom_walker::after_dom_children): Gate fma_deferring_state's enable predicate on param_widening_mul_defer_fma in addition to param_avoid_fma_max_bits. * config/aarch64/aarch64-tuning-flags.def (AVOID_CROSS_LOOP_FMA): Update the comment for the widening-mul-defer-fma effect. * config/aarch64/aarch64.cc (aarch64_override_options_internal): Inside the AARCH64_EXTRA_TUNE_AVOID_CROSS_LOOP_FMA block, set param_widening_mul_defer_fma to 0. gcc/testsuite/ChangeLog: * gcc.target/aarch64/widening-mul-defer-fma-1.c: New test. * gcc.target/aarch64/widening-mul-defer-fma-2.c: New test. * gcc.target/aarch64/widening-mul-defer-fma-3.c: New test. Signed-off-by: Philipp Tomsich --- Bootstrapped and regtested on aarch64-unknown-linux-gnu; no new regressions. The three new tests pass. Changes in v2: - Switch from a -fwidening-mul-defer-fma flag to a 0/1 --param=widening-mul-defer-fma, per review; drop the --param references from user-facing documentation. gcc/config/aarch64/aarch64-tuning-flags.def | 4 ++++ gcc/config/aarch64/aarch64.cc | 12 ++++++++--- gcc/params.opt | 6 ++++++ .../aarch64/widening-mul-defer-fma-1.c | 19 +++++++++++++++++ .../aarch64/widening-mul-defer-fma-2.c | 21 +++++++++++++++++++ .../aarch64/widening-mul-defer-fma-3.c | 21 +++++++++++++++++++ gcc/tree-ssa-math-opts.cc | 3 ++- 7 files changed, 82 insertions(+), 4 deletions(-) create mode 100644 gcc/testsuite/gcc.target/aarch64/widening-mul-defer-fma-1.c create mode 100644 gcc/testsuite/gcc.target/aarch64/widening-mul-defer-fma-2.c create mode 100644 gcc/testsuite/gcc.target/aarch64/widening-mul-defer-fma-3.c diff --git a/gcc/config/aarch64/aarch64-tuning-flags.def b/gcc/config/aarch64/aarch64-tuning-flags.def index 058dadecccaa..32bc1c4f3eec 100644 --- a/gcc/config/aarch64/aarch64-tuning-flags.def +++ b/gcc/config/aarch64/aarch64-tuning-flags.def @@ -40,6 +40,10 @@ AARCH64_EXTRA_TUNING_OPTION ("cse_sve_vl_constants", CSE_SVE_VL_CONSTANTS) AARCH64_EXTRA_TUNING_OPTION ("matched_vector_throughput", MATCHED_VECTOR_THROUGHPUT) +/* For cores whose pipeline disfavours loop-carried serial FMAs: set + avoid-fma-max-bits to its maximum (512, i.e. all FMA widths) to enable + the reassoc reorder for 3+ operand chains, and set widening-mul-defer-fma + to 0 to suppress the widening-mul pass's FMA deferring. */ AARCH64_EXTRA_TUNING_OPTION ("avoid_cross_loop_fma", AVOID_CROSS_LOOP_FMA) AARCH64_EXTRA_TUNING_OPTION ("fully_pipelined_fma", FULLY_PIPELINED_FMA) diff --git a/gcc/config/aarch64/aarch64.cc b/gcc/config/aarch64/aarch64.cc index 78f1eae8336c..1b1728e4e83d 100644 --- a/gcc/config/aarch64/aarch64.cc +++ b/gcc/config/aarch64/aarch64.cc @@ -20087,11 +20087,17 @@ aarch64_override_options_internal (struct gcc_options *opts, && opts->x_optimize >= aarch64_tune_params.prefetch->default_opt_level) opts->x_flag_prefetch_loop_arrays = 1; - /* Avoid loop-dependant FMA chains. */ + /* Avoid loop-dependant FMA chains. The reassoc-side reorder helper + keeps using --param=avoid-fma-max-bits; the widening-mul-side + deferring is gated separately by --param=widening-mul-defer-fma, so we + suppress only the deferring on these cores while leaving the reassoc + reorder active. */ if (aarch64_tune_params.extra_tuning_flags & AARCH64_EXTRA_TUNE_AVOID_CROSS_LOOP_FMA) - SET_OPTION_IF_UNSET (opts, opts_set, param_avoid_fma_max_bits, - 512); + { + SET_OPTION_IF_UNSET (opts, opts_set, param_avoid_fma_max_bits, 512); + SET_OPTION_IF_UNSET (opts, opts_set, param_widening_mul_defer_fma, 0); + } /* Consider fully pipelined FMA in reassociation. */ if (aarch64_tune_params.extra_tuning_flags diff --git a/gcc/params.opt b/gcc/params.opt index 90f9943c8cb4..044c4a10bc44 100644 --- a/gcc/params.opt +++ b/gcc/params.opt @@ -1323,4 +1323,10 @@ Maximum number of outgoing edges in a switch before VRP does not process it. Common Joined UInteger Var(param_vrp_vector_threshold) Init(250) Optimization Param Maximum number of basic blocks for VRP to use a basic cache vector. +-param=widening-mul-defer-fma= +Common Joined UInteger Var(param_widening_mul_defer_fma) Init(1) IntegerRange(0, 1) Param Optimization +When nonzero, the widening-multiply pass defers forming an FMA whose result +feeds a loop-header PHI, leaving a separate multiply and add. Set to zero to +contract such loop-carried reductions into an FMA instead. + ; This comment is to ensure we retain the blank line above. diff --git a/gcc/testsuite/gcc.target/aarch64/widening-mul-defer-fma-1.c b/gcc/testsuite/gcc.target/aarch64/widening-mul-defer-fma-1.c new file mode 100644 index 000000000000..dbee0ecb62b6 --- /dev/null +++ b/gcc/testsuite/gcc.target/aarch64/widening-mul-defer-fma-1.c @@ -0,0 +1,19 @@ +/* { dg-do compile } */ +/* { dg-options "-O2 -ffast-math -fno-tree-vectorize -mcpu=ampere1" } */ + +/* The ampere1 family sets AARCH64_EXTRA_TUNE_AVOID_CROSS_LOOP_FMA, which + sets widening-mul-defer-fma to 0 so that a 2-operand loop-carried + reduction is contracted to fmadd rather than deferred to fmul + fadd. + (avoid-fma-max-bits stays 512 to keep the reassoc reorder enabled.) */ + +double +dot (const double *a, const double *b, int n) +{ + double s = 0.0; + for (int i = 0; i < n; i++) + s += a[i] * b[i]; + return s; +} + +/* { dg-final { scan-assembler {\tfmadd\t} } } */ +/* { dg-final { scan-assembler-not {\tfmul\t} } } */ diff --git a/gcc/testsuite/gcc.target/aarch64/widening-mul-defer-fma-2.c b/gcc/testsuite/gcc.target/aarch64/widening-mul-defer-fma-2.c new file mode 100644 index 000000000000..a73d1ac1e9d3 --- /dev/null +++ b/gcc/testsuite/gcc.target/aarch64/widening-mul-defer-fma-2.c @@ -0,0 +1,21 @@ +/* { dg-do compile } */ +/* { dg-options "-O2 -ffast-math -fno-tree-vectorize --param=avoid-fma-max-bits=512" } */ + +/* With FMA deferring active (avoid-fma-max-bits > 0) and the default + widening-mul-defer-fma=1, the widening_mul pass refuses to form an FMA + whose result feeds the loop-header phi: the reduction stays as a + separate fmul + fadd. This is the behaviour the new param decouples + from the reassoc reorder (also gated by avoid-fma-max-bits). */ + +double +dot (const double *a, const double *b, int n) +{ + double s = 0.0; + for (int i = 0; i < n; i++) + s += a[i] * b[i]; + return s; +} + +/* { dg-final { scan-assembler {\tfmul\t} } } */ +/* { dg-final { scan-assembler {\tfadd\t} } } */ +/* { dg-final { scan-assembler-not {\tfmadd\t} } } */ diff --git a/gcc/testsuite/gcc.target/aarch64/widening-mul-defer-fma-3.c b/gcc/testsuite/gcc.target/aarch64/widening-mul-defer-fma-3.c new file mode 100644 index 000000000000..6463ca617734 --- /dev/null +++ b/gcc/testsuite/gcc.target/aarch64/widening-mul-defer-fma-3.c @@ -0,0 +1,21 @@ +/* { dg-do compile } */ +/* { dg-options "-O2 -ffast-math -fno-tree-vectorize --param=avoid-fma-max-bits=512 --param=widening-mul-defer-fma=0" } */ + +/* Companion to widening-mul-defer-fma-2.c: the same reduction with FMA + avoidance still requested for the reassoc reorder (avoid-fma-max-bits=512) + but --param=widening-mul-defer-fma=0 now suppresses the widening_mul + deferring, so the loop-carried multiply-add is contracted to a single + fmadd. This proves the new param decouples the deferring from + avoid-fma-max-bits. */ + +double +dot (const double *a, const double *b, int n) +{ + double s = 0.0; + for (int i = 0; i < n; i++) + s += a[i] * b[i]; + return s; +} + +/* { dg-final { scan-assembler {\tfmadd\t} } } */ +/* { dg-final { scan-assembler-not {\tfmul\t} } } */ diff --git a/gcc/tree-ssa-math-opts.cc b/gcc/tree-ssa-math-opts.cc index f0ede668d95e..c4a1d7bcf0bd 100644 --- a/gcc/tree-ssa-math-opts.cc +++ b/gcc/tree-ssa-math-opts.cc @@ -6607,7 +6607,8 @@ math_opts_dom_walker::after_dom_children (basic_block bb) { gimple_stmt_iterator gsi; - fma_deferring_state fma_state (param_avoid_fma_max_bits > 0); + fma_deferring_state fma_state (param_avoid_fma_max_bits > 0 + && param_widening_mul_defer_fma); for (gphi_iterator psi_next, psi = gsi_start_phis (bb); !gsi_end_p (psi); psi = psi_next)