| Message ID | 20260902145254.77832-4-ktkachov@nvidia.com |
|---|---|
| State | New |
| Headers |
Return-Path: <gcc-patches-bounces~patchwork=sourceware.org@gcc.gnu.org> X-Original-To: patchwork@sourceware.org Delivered-To: patchwork@sourceware.org Received: from vm01.sourceware.org (localhost [IPv6:::1]) by sourceware.org (Postfix) with ESMTP id 71A204B9DB4B for <patchwork@sourceware.org>; Wed, 2 Sep 2026 14:54:24 +0000 (GMT) DKIM-Filter: OpenDKIM Filter v2.11.0 sourceware.org 71A204B9DB4B Authentication-Results: sourceware.org; dkim=pass (2048-bit key, unprotected) header.d=Nvidia.com header.i=@Nvidia.com header.a=rsa-sha256 header.s=selector2 header.b=Ib597/Rv X-Original-To: gcc-patches@gcc.gnu.org Delivered-To: gcc-patches@gcc.gnu.org Received: from CH5PR02CU005.outbound.protection.outlook.com (mail-northcentralusazlp170120005.outbound.protection.outlook.com [IPv6:2a01:111:f403:c105::5]) by sourceware.org (Postfix) with ESMTPS id C7D954BA9020 for <gcc-patches@gcc.gnu.org>; Wed, 2 Sep 2026 14:53:42 +0000 (GMT) DMARC-Filter: OpenDMARC Filter v1.4.2 sourceware.org C7D954BA9020 Authentication-Results: sourceware.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com Authentication-Results: sourceware.org; spf=fail smtp.mailfrom=nvidia.com ARC-Filter: OpenARC Filter v1.0.0 sourceware.org C7D954BA9020 Authentication-Results: sourceware.org; arc=pass smtp.remote-ip=2a01:111:f403:c105::5 ARC-Seal: i=2; a=rsa-sha256; d=sourceware.org; s=key; t=1788360822; cv=pass; b=WcROGBJGSUcK2DmUmhly2QiZ50vO9xt+98aJPE6mCJO3L2dwzLmweL/AZi4IDLARJKhdhddGE5ncdSSODy8hcZzfKEOXRTIxiDA9NUSQHBBEEHtBT0uhF5Qm3kfbrjkLpT8FAXVqd9q0LDgQVxwpgtVn7H7SeiUDfuGoKDKK7yg= ARC-Message-Signature: i=2; a=rsa-sha256; d=sourceware.org; s=key; t=1788360822; c=relaxed/simple; bh=+e5m6poy2f4jemhqXGFLbSaUAK2+QtrCtTtEqcqBgGk=; h=DKIM-Signature:From:To:Subject:Date:Message-ID:MIME-Version; b=q/1s82ewaT3k605qgAJr5kpg3I+zvSDYGk0wPRg9Gw3W0plnAXa1RUrbu1VHEWaBZP5pU0LYdmBj1F3r6aE3BEHGWDVEfBbeNAfNqb62xhnG7NOYD1krKRjlF0Uy14fa77vTkpmPOoj7YOaReJ9sICyVkm8Zn78Lfr4eRt8M8jI= ARC-Authentication-Results: i=2; sourceware.org; dkim=pass (2048-bit key, unprotected) header.d=Nvidia.com header.i=@Nvidia.com header.a=rsa-sha256 header.s=selector2 header.b=Ib597/Rv DKIM-Filter: OpenDKIM Filter v2.11.0 sourceware.org C7D954BA9020 ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=lfA/VjGBxn4sOY0s7z7kJtV/KNCjtKzFml5hb+nxEzPp84feruot9q/btDmlXrJowwr4E+bz1e5j3uPy8RkE4wwxJmDymoL6Vt6LxHxXqbxl67dvLFl3G+W6avL3Tnoz0XS+ajWPZvNM5uRfn+Eb4oG3kFhOM+Ighc3PjHdSBnGuua+DTBV/hmA+VQlOzjxllc3pOcEdRiG1Grjbm52GrK3LsTOmDmy9hufl0i+gJel4WlizOKz29vk5iLhZ89uqy5WZUUx18Kl7TKUrdwMAR+Rid7cj8ztgCWD9KEHH0tIIS3ELnBjLyOS8f39SbPc76XiSIWXKttbLYExpFLXGvw== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=Brq6hIhwNbic0BPfyWkTL2o/YBoWUVmySl/sw6E6QGU=; b=m9W/frzhgCVN+wA6k4w6Mtt3WcoCzACWOBva8w7A7l+PtJkzulZLTqIqTQL4HCccciYKQa8NTaqm2y1MH7hS+l52th7jlTmYKlYYanFOgJFwKrwDWYmXJxrUIgz7F2vr0vTrF8kejeowKrsr4fkO0Saw4xiYILJzIlZuHO7ydxINpOHCGb8WSne1HkcffKvJlFu5XW76QQX0vxVAt7Z9/LUkhSfZoPytgFrExrcuuTV4DEDpdZP5b1u8fgCOBZn3KOmWLgvc9wftVwVPZ9bbxX8pkbDYqORqGaswebXjvvwUwBXlhv+CZ90AkVUmfFl4RGiOogmTwJ3ojyJ7emVerQ== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 216.228.117.160) smtp.rcpttodomain=gcc.gnu.org smtp.mailfrom=nvidia.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=nvidia.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=Nvidia.com; s=selector2; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=Brq6hIhwNbic0BPfyWkTL2o/YBoWUVmySl/sw6E6QGU=; b=Ib597/RvAzJ7RgukNPOQsJ218CBO3rSuRlSFVl9GW0eceZnUSp2pTe7jAEFB6745VwcUlt3vlVUSRi51enMcvPyawt6h3B+GJ54yLJTwzVtEzXOQeIKG01XbPa5kjjLVhTozBTQLRO2N6JBxGzFHLIBKYKumopke4LA6kbJAFamIeCRHnRCrJnv+qwS5/KMpntexiUr8vGecGZoWGR50PJozQ0TAE8pda1XcAY6jTkNtRGlVgLZ8h8ow/4BifulKhrFe8EZv1sygxLPJDxoPmB9IqArO3G+Zh04rB6gIyx2E8zi6CXWSS7DxzEslVPQto85t0UhwTvpDWQt8MjNU4g== Received: from BN9PR03CA0450.namprd03.prod.outlook.com (2603:10b6:408:113::35) by CH2PR12MB4261.namprd12.prod.outlook.com (2603:10b6:610:a9::8) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.360.13; Wed, 2 Sep 2026 14:53:37 +0000 Received: from BN3PEPF0000B06E.namprd21.prod.outlook.com (2603:10b6:408:113:cafe::5) by BN9PR03CA0450.outlook.office365.com (2603:10b6:408:113::35) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.382.10 via Frontend Transport; Wed, 2 Sep 2026 14:53:36 +0000 X-MS-Exchange-Authentication-Results: mx.microsoft.com 1; spf=pass (sender IP is 216.228.117.160) smtp.mailfrom=nvidia.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=nvidia.com; Received-SPF: Pass (protection.outlook.com: domain of nvidia.com designates 216.228.117.160 as permitted sender) receiver=protection.outlook.com; client-ip=216.228.117.160; helo=mail.nvidia.com; pr=C Received: from mail.nvidia.com (216.228.117.160) by BN3PEPF0000B06E.mail.protection.outlook.com (10.167.243.73) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.406.1 via Frontend Transport; Wed, 2 Sep 2026 14:53:36 +0000 Received: from rnnvmail201.nvidia.com (10.129.68.8) by mail.nvidia.com (10.129.200.66) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.46; Wed, 2 Sep 2026 07:53:09 -0700 Received: from ktkachov-mlt.nvidia.com (10.126.230.37) by rnnvmail201.nvidia.com (10.129.68.8) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.46; Wed, 2 Sep 2026 07:53:08 -0700 From: <ktkachov@nvidia.com> To: <gcc-patches@gcc.gnu.org> CC: Kyrylo Tkachov <ktkachov@nvidia.com> Subject: [PATCH 02/13] match-sat-alu.pd: Recognize X - MIN (X, Y) as saturating subtraction Date: Wed, 2 Sep 2026 16:52:43 +0200 Message-ID: <20260902145254.77832-4-ktkachov@nvidia.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260902145254.77832-1-ktkachov@nvidia.com> References: <20260902145254.77832-1-ktkachov@nvidia.com> MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Content-Type: text/plain X-Originating-IP: [10.126.230.37] X-ClientProxiedBy: rnnvmail202.nvidia.com (10.129.68.7) To rnnvmail201.nvidia.com (10.129.68.8) X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: BN3PEPF0000B06E:EE_|CH2PR12MB4261:EE_ X-MS-Office365-Filtering-Correlation-Id: ddb205a2-296b-4b82-e37f-08df0901f7f0 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0; ARA:13230040|23010399003|36860700016|82310400026|1800799024|376014|13003099007|6133799003|10067099003|56012099006|11063799006|3023799007|22082099003|18002099003; X-Microsoft-Antispam-Message-Info: E5aEqAmYV+ncjHyqHc1DBISP+IsyTahb4oHZmNzSfAjOre4I5nKcQRJD1ZKPCvQeTde7uxwre/WTf0qZzOH+/Cpq41cqU4ld2wjOdkX5x6Y4EdFzGXKQbco4CrnxzOj09wYt9TMNATGyHAcMj6B4ktyGKm0HNKznXmYUBxS0WhbPo70hwzHPCtTJ7BzQulXEXWsHujTPN9xSsAyHHrmNKabfUDaTwzUn53WAMcZtKdp5RnMMKZLtjRR0JtgHWpe5dvUS9/Yt5hD8+FfTPaswm5pm5NBVpg4+e/D+h+S6HBPX0FfcAIsQ//xkOUIXl2Zhm6ORSB+f3izf/p45mFwLTtpAFUODerGG4I4LakZ2uBDFgvvwBtP/a3Ad7CNM7xjlBMRwETALKBiZgCAn1y3ROcEUpnb0gK4cl0T55pWInoPxEB58tCzpHruSWP36Coft2jE8JP/NDPYz1MgPFvxlTfBAkwFiAHUbgpJNdHhhcTOJnUR4/9vbH81vGkn1PYjrmgyTTiSzaVDNDwkhbH+cbghl+VA8N9EylWhJroAoMERczjMygnUndiogsxL9lHAsvi6Agmyr5IGXh6/2YPSI6wmx8KsJnpMPo4c4nqXDzfUkKlDQ4X89WwHarNd3G+uECZAbupNqJ4r9dRnao7j72fA9P3q7OzhI940Wuft8iyAWF7mHWqnvO1MqnczZS7DFof2gOLYrzpldtigH8Gw+yA== X-Forefront-Antispam-Report: CIP:216.228.117.160; CTRY:US; LANG:en; SCL:1; SRV:; IPV:NLI; SFV:NSPM; H:mail.nvidia.com; PTR:dc6edge1.nvidia.com; CAT:NONE; SFS:(13230040)(23010399003)(36860700016)(82310400026)(1800799024)(376014)(13003099007)(6133799003)(10067099003)(56012099006)(11063799006)(3023799007)(22082099003)(18002099003); DIR:OUT; SFP:1101; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: q3cAapguC1an48nAy80nfJTjr2lNfaADhLLxn9SBItEOq63WZOxW6JWy596ZsBE04K+FxNjmbDbbk+7sJPWl6DajgjLawQGrrfpEKTSyefT/Z+9pvSHYPU7mDFVhJc2x0uEXj+92bxJ+MHQs41vFD5Nknj2xjUuTYJcmGBQWCRV2m9jKLKtiNGcoQoc9siVTpuBg7X2QTVNHDGL7Tey4n4T7g6WyI88ydyk8FzwMsg26tp9W7AMer1W8QJDYcrQfMwC/EQyI+8b3Z10dsVSXq23Ej14hPdV81Hrr0WJAd/YmFWSY20ilfwME866Bp1LHdM31YKL3xcC7cxNHZfU0aVOHs0riirnZI6TjBBVat8trR2kIO/tTv9dG5WmSbrbKUPql4isqU1MkczjEsgqB1elnMcyudaNKNYF3Z1rcwHgNNwWtX6B+tnI6EkLC+1Ep X-OriginatorOrg: Nvidia.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 02 Sep 2026 14:53:36.7447 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: ddb205a2-296b-4b82-e37f-08df0901f7f0 X-MS-Exchange-CrossTenant-Id: 43083d15-7273-40c1-b7db-39efd9ccc17a X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=43083d15-7273-40c1-b7db-39efd9ccc17a; Ip=[216.228.117.160]; Helo=[mail.nvidia.com] X-MS-Exchange-CrossTenant-AuthSource: BN3PEPF0000B06E.namprd21.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: CH2PR12MB4261 X-Spam-Status: No, score=-8.1 required=5.0 tests=BAYES_00, DKIMWL_WL_HIGH, DKIM_SIGNED, DKIM_VALID, DKIM_VALID_AU, DKIM_VALID_EF, FORGED_SPF_HELO, GIT_PATCH_0, KAM_SHORT, LOCAL_AUTHENTICATION_FAIL_SPF, RCVD_IN_DNSWL_NONE, SPF_HELO_PASS, SPF_NONE, TXREP shortcircuit=no autolearn=ham autolearn_force=no version=3.4.6 X-Spam-Checker-Version: SpamAssassin 3.4.6 (2021-04-09) on sourceware.org X-BeenThere: gcc-patches@gcc.gnu.org X-Mailman-Version: 2.1.30 Precedence: list List-Id: Gcc-patches mailing list <gcc-patches.gcc.gnu.org> List-Unsubscribe: <https://gcc.gnu.org/mailman/options/gcc-patches>, <mailto:gcc-patches-request@gcc.gnu.org?subject=unsubscribe> List-Archive: <https://gcc.gnu.org/pipermail/gcc-patches/> List-Post: <mailto:gcc-patches@gcc.gnu.org> List-Help: <mailto:gcc-patches-request@gcc.gnu.org?subject=help> List-Subscribe: <https://gcc.gnu.org/mailman/listinfo/gcc-patches>, <mailto:gcc-patches-request@gcc.gnu.org?subject=subscribe> Errors-To: gcc-patches-bounces~patchwork=sourceware.org@gcc.gnu.org |
| Series |
Saturating arithmetic matching improvements
|
|
Commit Message
Kyrylo Tkachov
Sept. 2, 2026, 2:52 p.m. UTC
From: Kyrylo Tkachov <ktkachov@nvidia.com> X - MIN (X, Y) is unsigned saturating subtraction, but subtraction statements were not checked for this form. Recognize the form and check MINUS_EXPR statements for saturation. Require the MIN to have one use so that the replacement removes it. If a replacement changes the current statement, do not try to widen that statement. AArch64 -O2: before: cmp w1, w0 csel w1, w1, w0, ls sub w0, w0, w1 after: subs w0, w0, w1 csel w0, w0, wzr, cs The vector form becomes one UQSUB instead of a comparison, a select, and a subtraction. Bootstrapped and tested on aarch64-none-linux-gnu. Ok for trunk? gcc/ChangeLog: * match-sat-alu.pd (unsigned_integer_sat_sub): Add the X - MIN (X, Y) form. * tree-ssa-math-opts.cc (math_opts_dom_walker::before_dom_children): Call match_unsigned_saturation_sub for MINUS_EXPR as well, and only widen while the statement is unchanged. gcc/testsuite/ChangeLog: * gcc.target/aarch64/sat_u_sub_minmax-1.c: New test. * gcc.target/aarch64/sat_u_sub_minmax-2.c: New test. * gcc.target/aarch64/sat_u_sub_minmax-3.c: New test. Signed-off-by: Kyrylo Tkachov <ktkachov@nvidia.com> --- gcc/match-sat-alu.pd | 6 ++++ .../gcc.target/aarch64/sat_u_sub_minmax-1.c | 18 +++++++++++ .../gcc.target/aarch64/sat_u_sub_minmax-2.c | 32 +++++++++++++++++++ .../gcc.target/aarch64/sat_u_sub_minmax-3.c | 27 ++++++++++++++++ gcc/tree-ssa-math-opts.cc | 5 +-- 5 files changed, 86 insertions(+), 2 deletions(-) create mode 100644 gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-1.c create mode 100644 gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-2.c create mode 100644 gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-3.c
Comments
On Wed, Sep 2, 2026 at 7:54 AM <ktkachov@nvidia.com> wrote: > > From: Kyrylo Tkachov <ktkachov@nvidia.com> > > X - MIN (X, Y) is unsigned saturating subtraction, but subtraction > statements were not checked for this form. > > Recognize the form and check MINUS_EXPR statements for saturation. Require > the MIN to have one use so that the replacement removes it. If a replacement > changes the current statement, do not try to widen that statement. > > AArch64 -O2: > > before: > > cmp w1, w0 > csel w1, w1, w0, ls > sub w0, w0, w1 > > after: > > subs w0, w0, w1 > csel w0, w0, wzr, cs > > The vector form becomes one UQSUB instead of a comparison, a select, and a > subtraction. > > Bootstrapped and tested on aarch64-none-linux-gnu. > > Ok for trunk? > > gcc/ChangeLog: > > * match-sat-alu.pd (unsigned_integer_sat_sub): Add the > X - MIN (X, Y) form. > * tree-ssa-math-opts.cc > (math_opts_dom_walker::before_dom_children): Call > match_unsigned_saturation_sub for MINUS_EXPR as well, and > only widen while the statement is unchanged. > > gcc/testsuite/ChangeLog: > > * gcc.target/aarch64/sat_u_sub_minmax-1.c: New test. > * gcc.target/aarch64/sat_u_sub_minmax-2.c: New test. > * gcc.target/aarch64/sat_u_sub_minmax-3.c: New test. > > Signed-off-by: Kyrylo Tkachov <ktkachov@nvidia.com> > --- > gcc/match-sat-alu.pd | 6 ++++ > .../gcc.target/aarch64/sat_u_sub_minmax-1.c | 18 +++++++++++ > .../gcc.target/aarch64/sat_u_sub_minmax-2.c | 32 +++++++++++++++++++ > .../gcc.target/aarch64/sat_u_sub_minmax-3.c | 27 ++++++++++++++++ > gcc/tree-ssa-math-opts.cc | 5 +-- > 5 files changed, 86 insertions(+), 2 deletions(-) > create mode 100644 gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-1.c > create mode 100644 gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-2.c > create mode 100644 gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-3.c > > diff --git a/gcc/match-sat-alu.pd b/gcc/match-sat-alu.pd > index 7156529f589..1b35b35cff8 100644 > --- a/gcc/match-sat-alu.pd > +++ b/gcc/match-sat-alu.pd > @@ -131,6 +131,12 @@ along with GCC; see the file COPYING3. If not see > /* SAT_U_SUB = (X - Y) * (X >= Y) */ > (mult:c (minus @0 @1) (convert (ge @0 @1))) > (if (types_match (type, @0, @1)))) > + (match (unsigned_integer_sat_sub @0 @1) > + /* SAT_U_SUB = X - MIN (X, Y). The MIN has to be single use: while it > + stays live the saturating subtract is computed beside it instead of > + replacing it, which costs an instruction. */ > + (minus @0 (min:c@2 @0 @1)) > + (if (single_use (@2)))) :s on the min. So `(minus @0 (min:cs @0 @1))` > (match (unsigned_integer_sat_sub @0 @1) > /* DIFF = SUB_OVERFLOW (X, Y) > SAT_U_SUB = REALPART (DIFF) | (IMAGPART (DIFF) + (-1)) */ > diff --git a/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-1.c b/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-1.c > new file mode 100644 > index 00000000000..aa1d7b6f031 > --- /dev/null > +++ b/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-1.c > @@ -0,0 +1,18 @@ > +/* { dg-do compile } */ > +/* { dg-options "-O2 -fdump-tree-optimized" } */ > + > +#define DEF_MIN(T) \ > + T min_##T (T a, T b) { return a - (a < b ? a : b); } \ > + T nim_##T (T a, T b) { return a - (b < a ? b : a); } > + > +typedef unsigned char u8; > +typedef unsigned short u16; > +typedef unsigned int u32; > +typedef unsigned long long u64; > + > +DEF_MIN (u8) > +DEF_MIN (u16) > +DEF_MIN (u32) > +DEF_MIN (u64) > + > +/* { dg-final { scan-tree-dump-times "\\.SAT_SUB " 8 "optimized" } } */ > diff --git a/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-2.c b/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-2.c > new file mode 100644 > index 00000000000..dcf8ab46309 > --- /dev/null > +++ b/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-2.c > @@ -0,0 +1,32 @@ > +/* { dg-do compile } */ > +/* { dg-options "-O3 -fdump-tree-optimized" } */ > + > +typedef unsigned char u8; > +typedef unsigned int u32; > + > +/* The recognition must reach the vectoriser, so that the loop becomes a > + single uqsub rather than a compare, a select and a subtraction. */ > + > +void > +min_loop (u8 *__restrict d, u8 *__restrict a, u8 *__restrict b, int n) > +{ > + for (int i = 0; i < n; i++) > + { > + u8 x = a[i], y = b[i]; > + d[i] = x - (x < y ? x : y); > + } > +} > + > +void > +min_loop32 (u32 *__restrict d, u32 *__restrict a, u32 *__restrict b, int n) > +{ > + for (int i = 0; i < n; i++) > + { > + u32 x = a[i], y = b[i]; > + d[i] = x - (y < x ? y : x); > + } > +} > + > +/* { dg-final { scan-tree-dump "\\.SAT_SUB " "optimized" } } */ > +/* { dg-final { scan-assembler "uqsub\tv\[0-9\]+\\.16b" } } */ > +/* { dg-final { scan-assembler "uqsub\tv\[0-9\]+\\.4s" } } */ > diff --git a/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-3.c b/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-3.c > new file mode 100644 > index 00000000000..029c136b7d0 > --- /dev/null > +++ b/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-3.c > @@ -0,0 +1,27 @@ > +/* { dg-do compile } */ > +/* { dg-options "-O2 -fdump-tree-optimized" } */ > + > +typedef unsigned int u32; > + > +/* While the MIN stays live the saturating subtract would be computed beside > + it rather than instead of it, so the rules do not fire. */ > + > +u32 > +min_live (u32 a, u32 b, u32 *o) > +{ > + u32 m = a < b ? a : b; > + *o = m; > + return a - m; > +} > + > +u32 > +add_min_live (u32 a, u32 b, u32 *o) > +{ > + u32 t = ~a; > + u32 m = b < t ? b : t; > + *o = m; > + return a + m; > +} > + > +/* { dg-final { scan-tree-dump-not "\\.SAT_SUB " "optimized" } } */ > +/* { dg-final { scan-tree-dump-not "\\.SAT_ADD " "optimized" } } */ > diff --git a/gcc/tree-ssa-math-opts.cc b/gcc/tree-ssa-math-opts.cc > index b371b5b7cff..ed3abb7d5c4 100644 > --- a/gcc/tree-ssa-math-opts.cc > +++ b/gcc/tree-ssa-math-opts.cc > @@ -7332,10 +7332,11 @@ math_opts_dom_walker::after_dom_children (basic_block bb) > > case PLUS_EXPR: > match_saturation_add_with_assign (&gsi, as_a<gassign *> (stmt)); > - match_unsigned_saturation_sub (&gsi, as_a<gassign *> (stmt)); > /* fall-through */ > case MINUS_EXPR: > - if (!convert_plusminus_to_widen (&gsi, stmt, code)) > + match_unsigned_saturation_sub (&gsi, as_a<gassign *> (stmt)); > + if (gsi_stmt (gsi) == stmt > + && !convert_plusminus_to_widen (&gsi, stmt, code)) Instead of `gsi_stmt (gsi) == stmt` instead return true from match_unsigned_saturation_sub if something was done. And do: if (!match_unsigned_saturation_sub (&gsi, as_a<gassign *> (stmt)) && !convert_plusminus_to_widen (&gsi, stmt, code)) { ... > { > match_arith_overflow (&gsi, stmt, code, m_cfg_changed_p); > if (gsi_stmt (gsi) == stmt) > -- > 2.50.1 (Apple Git-155) >
> On 3 Sep 2026, at 08:00, Andrea Pinski <andrew.pinski@oss.qualcomm.com> wrote: > > On Wed, Sep 2, 2026 at 7:54 AM <ktkachov@nvidia.com> wrote: >> >> From: Kyrylo Tkachov <ktkachov@nvidia.com> >> >> X - MIN (X, Y) is unsigned saturating subtraction, but subtraction >> statements were not checked for this form. >> >> Recognize the form and check MINUS_EXPR statements for saturation. Require >> the MIN to have one use so that the replacement removes it. If a replacement >> changes the current statement, do not try to widen that statement. >> >> AArch64 -O2: >> >> before: >> >> cmp w1, w0 >> csel w1, w1, w0, ls >> sub w0, w0, w1 >> >> after: >> >> subs w0, w0, w1 >> csel w0, w0, wzr, cs >> >> The vector form becomes one UQSUB instead of a comparison, a select, and a >> subtraction. >> >> Bootstrapped and tested on aarch64-none-linux-gnu. >> >> Ok for trunk? >> >> gcc/ChangeLog: >> >> * match-sat-alu.pd (unsigned_integer_sat_sub): Add the >> X - MIN (X, Y) form. >> * tree-ssa-math-opts.cc >> (math_opts_dom_walker::before_dom_children): Call >> match_unsigned_saturation_sub for MINUS_EXPR as well, and >> only widen while the statement is unchanged. >> >> gcc/testsuite/ChangeLog: >> >> * gcc.target/aarch64/sat_u_sub_minmax-1.c: New test. >> * gcc.target/aarch64/sat_u_sub_minmax-2.c: New test. >> * gcc.target/aarch64/sat_u_sub_minmax-3.c: New test. >> >> Signed-off-by: Kyrylo Tkachov <ktkachov@nvidia.com> >> --- >> gcc/match-sat-alu.pd | 6 ++++ >> .../gcc.target/aarch64/sat_u_sub_minmax-1.c | 18 +++++++++++ >> .../gcc.target/aarch64/sat_u_sub_minmax-2.c | 32 +++++++++++++++++++ >> .../gcc.target/aarch64/sat_u_sub_minmax-3.c | 27 ++++++++++++++++ >> gcc/tree-ssa-math-opts.cc | 5 +-- >> 5 files changed, 86 insertions(+), 2 deletions(-) >> create mode 100644 gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-1.c >> create mode 100644 gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-2.c >> create mode 100644 gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-3.c >> >> diff --git a/gcc/match-sat-alu.pd b/gcc/match-sat-alu.pd >> index 7156529f589..1b35b35cff8 100644 >> --- a/gcc/match-sat-alu.pd >> +++ b/gcc/match-sat-alu.pd >> @@ -131,6 +131,12 @@ along with GCC; see the file COPYING3. If not see >> /* SAT_U_SUB = (X - Y) * (X >= Y) */ >> (mult:c (minus @0 @1) (convert (ge @0 @1))) >> (if (types_match (type, @0, @1)))) >> + (match (unsigned_integer_sat_sub @0 @1) >> + /* SAT_U_SUB = X - MIN (X, Y). The MIN has to be single use: while it >> + stays live the saturating subtract is computed beside it instead of >> + replacing it, which costs an instruction. */ >> + (minus @0 (min:c@2 @0 @1)) >> + (if (single_use (@2)))) > > :s on the min. > So `(minus @0 (min:cs @0 @1))` The :s does not enforce single_use in a named predicate matcher. I haven’t looked into whether that’s by design or an oversight though. Is that something that’s implementable? > >> (match (unsigned_integer_sat_sub @0 @1) >> /* DIFF = SUB_OVERFLOW (X, Y) >> SAT_U_SUB = REALPART (DIFF) | (IMAGPART (DIFF) + (-1)) */ >> diff --git a/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-1.c b/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-1.c >> new file mode 100644 >> index 00000000000..aa1d7b6f031 >> --- /dev/null >> +++ b/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-1.c >> @@ -0,0 +1,18 @@ >> +/* { dg-do compile } */ >> +/* { dg-options "-O2 -fdump-tree-optimized" } */ >> + >> +#define DEF_MIN(T) \ >> + T min_##T (T a, T b) { return a - (a < b ? a : b); } \ >> + T nim_##T (T a, T b) { return a - (b < a ? b : a); } >> + >> +typedef unsigned char u8; >> +typedef unsigned short u16; >> +typedef unsigned int u32; >> +typedef unsigned long long u64; >> + >> +DEF_MIN (u8) >> +DEF_MIN (u16) >> +DEF_MIN (u32) >> +DEF_MIN (u64) >> + >> +/* { dg-final { scan-tree-dump-times "\\.SAT_SUB " 8 "optimized" } } */ >> diff --git a/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-2.c b/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-2.c >> new file mode 100644 >> index 00000000000..dcf8ab46309 >> --- /dev/null >> +++ b/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-2.c >> @@ -0,0 +1,32 @@ >> +/* { dg-do compile } */ >> +/* { dg-options "-O3 -fdump-tree-optimized" } */ >> + >> +typedef unsigned char u8; >> +typedef unsigned int u32; >> + >> +/* The recognition must reach the vectoriser, so that the loop becomes a >> + single uqsub rather than a compare, a select and a subtraction. */ >> + >> +void >> +min_loop (u8 *__restrict d, u8 *__restrict a, u8 *__restrict b, int n) >> +{ >> + for (int i = 0; i < n; i++) >> + { >> + u8 x = a[i], y = b[i]; >> + d[i] = x - (x < y ? x : y); >> + } >> +} >> + >> +void >> +min_loop32 (u32 *__restrict d, u32 *__restrict a, u32 *__restrict b, int n) >> +{ >> + for (int i = 0; i < n; i++) >> + { >> + u32 x = a[i], y = b[i]; >> + d[i] = x - (y < x ? y : x); >> + } >> +} >> + >> +/* { dg-final { scan-tree-dump "\\.SAT_SUB " "optimized" } } */ >> +/* { dg-final { scan-assembler "uqsub\tv\[0-9\]+\\.16b" } } */ >> +/* { dg-final { scan-assembler "uqsub\tv\[0-9\]+\\.4s" } } */ >> diff --git a/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-3.c b/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-3.c >> new file mode 100644 >> index 00000000000..029c136b7d0 >> --- /dev/null >> +++ b/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-3.c >> @@ -0,0 +1,27 @@ >> +/* { dg-do compile } */ >> +/* { dg-options "-O2 -fdump-tree-optimized" } */ >> + >> +typedef unsigned int u32; >> + >> +/* While the MIN stays live the saturating subtract would be computed beside >> + it rather than instead of it, so the rules do not fire. */ >> + >> +u32 >> +min_live (u32 a, u32 b, u32 *o) >> +{ >> + u32 m = a < b ? a : b; >> + *o = m; >> + return a - m; >> +} >> + >> +u32 >> +add_min_live (u32 a, u32 b, u32 *o) >> +{ >> + u32 t = ~a; >> + u32 m = b < t ? b : t; >> + *o = m; >> + return a + m; >> +} >> + >> +/* { dg-final { scan-tree-dump-not "\\.SAT_SUB " "optimized" } } */ >> +/* { dg-final { scan-tree-dump-not "\\.SAT_ADD " "optimized" } } */ >> diff --git a/gcc/tree-ssa-math-opts.cc b/gcc/tree-ssa-math-opts.cc >> index b371b5b7cff..ed3abb7d5c4 100644 >> --- a/gcc/tree-ssa-math-opts.cc >> +++ b/gcc/tree-ssa-math-opts.cc >> @@ -7332,10 +7332,11 @@ math_opts_dom_walker::after_dom_children (basic_block bb) >> >> case PLUS_EXPR: >> match_saturation_add_with_assign (&gsi, as_a<gassign *> (stmt)); >> - match_unsigned_saturation_sub (&gsi, as_a<gassign *> (stmt)); >> /* fall-through */ >> case MINUS_EXPR: >> - if (!convert_plusminus_to_widen (&gsi, stmt, code)) >> + match_unsigned_saturation_sub (&gsi, as_a<gassign *> (stmt)); >> + if (gsi_stmt (gsi) == stmt >> + && !convert_plusminus_to_widen (&gsi, stmt, code)) > > Instead of `gsi_stmt (gsi) == stmt` instead return true from > match_unsigned_saturation_sub if something was done. > And do: > if (!match_unsigned_saturation_sub (&gsi, as_a<gassign *> (stmt)) > && !convert_plusminus_to_widen (&gsi, stmt, code)) > { > ... Thanks, I’ll do that in the v2. Kyrill > >> { >> match_arith_overflow (&gsi, stmt, code, m_cfg_changed_p); >> if (gsi_stmt (gsi) == stmt) >> -- >> 2.50.1 (Apple Git-155)
diff --git a/gcc/match-sat-alu.pd b/gcc/match-sat-alu.pd index 7156529f589..1b35b35cff8 100644 --- a/gcc/match-sat-alu.pd +++ b/gcc/match-sat-alu.pd @@ -131,6 +131,12 @@ along with GCC; see the file COPYING3. If not see /* SAT_U_SUB = (X - Y) * (X >= Y) */ (mult:c (minus @0 @1) (convert (ge @0 @1))) (if (types_match (type, @0, @1)))) + (match (unsigned_integer_sat_sub @0 @1) + /* SAT_U_SUB = X - MIN (X, Y). The MIN has to be single use: while it + stays live the saturating subtract is computed beside it instead of + replacing it, which costs an instruction. */ + (minus @0 (min:c@2 @0 @1)) + (if (single_use (@2)))) (match (unsigned_integer_sat_sub @0 @1) /* DIFF = SUB_OVERFLOW (X, Y) SAT_U_SUB = REALPART (DIFF) | (IMAGPART (DIFF) + (-1)) */ diff --git a/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-1.c b/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-1.c new file mode 100644 index 00000000000..aa1d7b6f031 --- /dev/null +++ b/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-1.c @@ -0,0 +1,18 @@ +/* { dg-do compile } */ +/* { dg-options "-O2 -fdump-tree-optimized" } */ + +#define DEF_MIN(T) \ + T min_##T (T a, T b) { return a - (a < b ? a : b); } \ + T nim_##T (T a, T b) { return a - (b < a ? b : a); } + +typedef unsigned char u8; +typedef unsigned short u16; +typedef unsigned int u32; +typedef unsigned long long u64; + +DEF_MIN (u8) +DEF_MIN (u16) +DEF_MIN (u32) +DEF_MIN (u64) + +/* { dg-final { scan-tree-dump-times "\\.SAT_SUB " 8 "optimized" } } */ diff --git a/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-2.c b/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-2.c new file mode 100644 index 00000000000..dcf8ab46309 --- /dev/null +++ b/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-2.c @@ -0,0 +1,32 @@ +/* { dg-do compile } */ +/* { dg-options "-O3 -fdump-tree-optimized" } */ + +typedef unsigned char u8; +typedef unsigned int u32; + +/* The recognition must reach the vectoriser, so that the loop becomes a + single uqsub rather than a compare, a select and a subtraction. */ + +void +min_loop (u8 *__restrict d, u8 *__restrict a, u8 *__restrict b, int n) +{ + for (int i = 0; i < n; i++) + { + u8 x = a[i], y = b[i]; + d[i] = x - (x < y ? x : y); + } +} + +void +min_loop32 (u32 *__restrict d, u32 *__restrict a, u32 *__restrict b, int n) +{ + for (int i = 0; i < n; i++) + { + u32 x = a[i], y = b[i]; + d[i] = x - (y < x ? y : x); + } +} + +/* { dg-final { scan-tree-dump "\\.SAT_SUB " "optimized" } } */ +/* { dg-final { scan-assembler "uqsub\tv\[0-9\]+\\.16b" } } */ +/* { dg-final { scan-assembler "uqsub\tv\[0-9\]+\\.4s" } } */ diff --git a/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-3.c b/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-3.c new file mode 100644 index 00000000000..029c136b7d0 --- /dev/null +++ b/gcc/testsuite/gcc.target/aarch64/sat_u_sub_minmax-3.c @@ -0,0 +1,27 @@ +/* { dg-do compile } */ +/* { dg-options "-O2 -fdump-tree-optimized" } */ + +typedef unsigned int u32; + +/* While the MIN stays live the saturating subtract would be computed beside + it rather than instead of it, so the rules do not fire. */ + +u32 +min_live (u32 a, u32 b, u32 *o) +{ + u32 m = a < b ? a : b; + *o = m; + return a - m; +} + +u32 +add_min_live (u32 a, u32 b, u32 *o) +{ + u32 t = ~a; + u32 m = b < t ? b : t; + *o = m; + return a + m; +} + +/* { dg-final { scan-tree-dump-not "\\.SAT_SUB " "optimized" } } */ +/* { dg-final { scan-tree-dump-not "\\.SAT_ADD " "optimized" } } */ diff --git a/gcc/tree-ssa-math-opts.cc b/gcc/tree-ssa-math-opts.cc index b371b5b7cff..ed3abb7d5c4 100644 --- a/gcc/tree-ssa-math-opts.cc +++ b/gcc/tree-ssa-math-opts.cc @@ -7332,10 +7332,11 @@ math_opts_dom_walker::after_dom_children (basic_block bb) case PLUS_EXPR: match_saturation_add_with_assign (&gsi, as_a<gassign *> (stmt)); - match_unsigned_saturation_sub (&gsi, as_a<gassign *> (stmt)); /* fall-through */ case MINUS_EXPR: - if (!convert_plusminus_to_widen (&gsi, stmt, code)) + match_unsigned_saturation_sub (&gsi, as_a<gassign *> (stmt)); + if (gsi_stmt (gsi) == stmt + && !convert_plusminus_to_widen (&gsi, stmt, code)) { match_arith_overflow (&gsi, stmt, code, m_cfg_changed_p); if (gsi_stmt (gsi) == stmt)