From patchwork Wed Jul 8 06:04:36 2026 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Kyrylo Tkachov X-Patchwork-Id: 138709 Return-Path: X-Original-To: patchwork@sourceware.org Delivered-To: patchwork@sourceware.org Received: from vm01.sourceware.org (localhost [IPv6:::1]) by sourceware.org (Postfix) with ESMTP id AB4B14BA2E11 for ; Wed, 8 Jul 2026 06:06:28 +0000 (GMT) DKIM-Filter: OpenDKIM Filter v2.11.0 sourceware.org AB4B14BA2E11 Authentication-Results: sourceware.org; dkim=pass (2048-bit key, unprotected) header.d=Nvidia.com header.i=@Nvidia.com header.a=rsa-sha256 header.s=selector2 header.b=aYvc3KPn X-Original-To: gcc-patches@gcc.gnu.org Delivered-To: gcc-patches@gcc.gnu.org Received: from DM1PR04CU001.outbound.protection.outlook.com (mail-centralusazon11010039.outbound.protection.outlook.com [52.101.61.39]) by sourceware.org (Postfix) with ESMTPS id 71EAD4BA2E24 for ; Wed, 8 Jul 2026 06:05:20 +0000 (GMT) DMARC-Filter: OpenDMARC Filter v1.4.2 sourceware.org 71EAD4BA2E24 Authentication-Results: sourceware.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com Authentication-Results: sourceware.org; spf=fail smtp.mailfrom=nvidia.com ARC-Filter: OpenARC Filter v1.0.0 sourceware.org 71EAD4BA2E24 Authentication-Results: sourceware.org; arc=pass smtp.remote-ip=52.101.61.39 ARC-Seal: i=2; a=rsa-sha256; d=sourceware.org; s=key; t=1783490726; cv=pass; b=W3fsCzCpkqshZZgJT/Z7GtphQ34meexWNGP6n9LufyXac0GNnmwMqnTw2qIH9dJZy+kgAGDxw5eWXj1Dc7d8GhszfCSz/aGy4yu4r5cLiLKQUn1EbxAsbtaZB1Qwlea3WRfoAxv1Mj9+adOXvpgRIGvHpblE7fIP8zAzohBQGrU= ARC-Message-Signature: i=2; a=rsa-sha256; d=sourceware.org; s=key; t=1783490726; c=relaxed/simple; bh=IGeOkoINzj/xnHXYW3/UKMX+ZHii93XOraA1BEJ4JXg=; h=DKIM-Signature:From:To:Subject:Date:Message-ID:MIME-Version; b=tqHuVld+Nz4oVPOMaXSSEKAYLXo1bif3zIWfaAKxTlyyWh2bZKyPXmV6Ee+oDznUPu+E35/mMjlsrqMnLBfkE7SCaiWOpWI4VZbkYe/TrGb9CG22YcbvA9wTZpmR1vJPy/71m90Y6Y6q6+nJMV8RKoaaHOTRJCQ6JTP9YE85QQ8= ARC-Authentication-Results: i=2; sourceware.org; dkim=pass (2048-bit key, unprotected) header.d=Nvidia.com header.i=@Nvidia.com header.a=rsa-sha256 header.s=selector2 header.b=aYvc3KPn DKIM-Filter: OpenDKIM Filter v2.11.0 sourceware.org 71EAD4BA2E24 ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=LtBFWZSCpStWxcFMgvQdW1Bv2T/ePvTM0elrU9p1yVKuB78wIIpehPABkkcVJn4a9Yhw9LYDbdjBkbhoAHQGmPn3K3rbSURvrLh6WqbeJfX24Rid/44y7jTz/7T244cZwqAfvbsqmLFQaUSq5B4H7faenw6fCf44C7StizeSfLRMdaSy8P5u8fpZ8EkzC03ssRyv34nSjIVl5DGMkt/kouxx3BqtYOULtrXhopnNqw2nbLeswH1ZIMkHL5qmQiRPnSxemIQizQJ20vTKheUw+gOrGylG2xQPYvvD5PhhH28Q1DQJy0PXe9GoR5iGsG55Ylq7i55T3Cvyu+5UafC2cA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=XnOxwreYxSVKUAA4lUqL3NFpedZNrgH4t+92OThVK9Q=; b=Y+a4OiaSPGg8vWAIknWCgXeLHFClnvVcre4ApLPqMs2p8F74xyFRrMgWwtVY19nUr8nJa/P3+BbFzRsL4BYrDuO5pOi0seVEj/9OOuVF2QuP+CS+Tk+ZdeoUSdiw3hDWUYrgo4wDmHJy2GG8RPuKKCQKU1K0PlWQE58lyZFBBi55Xbe6cXColqa+NmFjS1aVFbd5NfwmrYMmcopCnhPNmiPB7mGvkdyD0Ddo2xnehiPFZ8HsWsRD0ev8RvkQIb+uqbkwZALk1a4BWhGja2rBzYLHa/MsL9siWv7j2Ptw16AweD4lZYBjw+NK6jgsZTM1JeejzfUjcs7EQBZ9XqelzA== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 216.228.117.160) smtp.rcpttodomain=gcc.gnu.org smtp.mailfrom=nvidia.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=nvidia.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=Nvidia.com; s=selector2; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=XnOxwreYxSVKUAA4lUqL3NFpedZNrgH4t+92OThVK9Q=; b=aYvc3KPnZ9unqpQsw8z00DC1CIK3ofOP8PW4Bu8QndFSElNYqNCd5pn/7YqXEdRVrlbnjcp8FrFuZpQPlDZlzAVbpSCSGCD1HZQamm2OcWVL4FkNUfLBEbhJDu070TNVSRiJPHqgZAOFbrIzqZJZCMU0qxrkuD0m+Slu/LeRgIUe7z8bZuGMQLnmYSU4ydrQa41+GyRQWouF+WG9GHlqzdLR3k0xxTSsZn79WQnEtpcKxD3S12/7GdMxdRKbkaWYZOls1S6dUEXx2sjwiBcnF7+s4QXF/3X91clawxfMnScn2a1q3oEJmI4/ADBYSaMcFZi9C3ic/HyST9LSxSIdJA== Received: from DS1P221CA0007.NAMP221.PROD.OUTLOOK.COM (2603:10b6:8:451::18) by SJ1PR12MB6100.namprd12.prod.outlook.com (2603:10b6:a03:45d::20) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.181.10; Wed, 8 Jul 2026 06:05:13 +0000 Received: from DS1PEPF00017095.namprd03.prod.outlook.com (2603:10b6:8:451:cafe::2c) by DS1P221CA0007.outlook.office365.com (2603:10b6:8:451::18) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.202.10 via Frontend Transport; Wed, 8 Jul 2026 06:05:12 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 216.228.117.160) smtp.mailfrom=nvidia.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=nvidia.com; Received-SPF: Pass (protection.outlook.com: domain of nvidia.com designates 216.228.117.160 as permitted sender) receiver=protection.outlook.com; client-ip=216.228.117.160; helo=mail.nvidia.com; pr=C Received: from mail.nvidia.com (216.228.117.160) by DS1PEPF00017095.mail.protection.outlook.com (10.167.17.138) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.181.6 via Frontend Transport; Wed, 8 Jul 2026 06:05:12 +0000 Received: from rnnvmail201.nvidia.com (10.129.68.8) by mail.nvidia.com (10.129.200.66) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.20; Tue, 7 Jul 2026 23:04:54 -0700 Received: from ktkachov-mlt.nvidia.com (10.126.231.37) by rnnvmail201.nvidia.com (10.129.68.8) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.20; Tue, 7 Jul 2026 23:04:52 -0700 From: To: CC: , , Kyrylo Tkachov Subject: [PATCH 1/2] tree-optimization/126043 - form comparison masks in early phiopt Date: Wed, 8 Jul 2026 08:04:36 +0200 Message-ID: <20260708060437.26491-1-ktkachov@nvidia.com> X-Mailer: git-send-email 2.50.1 MIME-Version: 1.0 X-Originating-IP: [10.126.231.37] X-ClientProxiedBy: rnnvmail202.nvidia.com (10.129.68.7) To rnnvmail201.nvidia.com (10.129.68.8) X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: DS1PEPF00017095:EE_|SJ1PR12MB6100:EE_ X-MS-Office365-Filtering-Correlation-Id: b8559cb8-10f7-4a87-bf03-08dedcb6dfa9 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0; ARA:13230040|376014|82310400026|36860700016|1800799024|23010399003|5023799004|11063799006|56012099006|3023799007|18002099003|6133799003|13003099007; X-Microsoft-Antispam-Message-Info: /eFYpBlJyxVLVF4INtR7ZT98Phajs//+sYV6ICNl8mHz1ZS0o09xyBmcXI+Dpr/XJdCpFN5PTuk5p6yGMZBjZgC4mcKZMlfvJtzFNTCf0fCws61NCcSNN1IQZDmgX8Kp6XZejj4mBGkWoddgpXeWxrc094OiaW22jmOyngqfzA+g8ZWre3doN+5ouBWpvUiFFWWPbcMdqyanwLwWbPMeGAB/yUOp91e7suDV/h7wEAcDOlYvRS9Jzpg6aw13Xt4wDxFG23JhAVcru/oJtPIfDMrRy0tcyd79CCi+vqgz5KKCAScmp9h5XOc8IOunWcYlSsTWeBAMB3Ra6rnOIvY5ZaHT/O3/6G93lNoUd1uhz8XipZupWNWyVqw+y/QRQBEkC2T24QNfklGO7Wd5V2Msdv9KTjtPezFCunElJljBWldUeN4OFNs4wrwOXCk8hgFMWcDSprLQBqKi3Bnx4ZyXbaIZ2nCNgmOLGkJKuSq07y3iFZNgR2BeIqEQSwiwxg6i4Y9gyylqHIA1Nauiio1elneeUvgxrlLxH+/sEIsNHdpq2fKD0X8IrUkIekr5RM6wyd0RYOIMYJAsqmLUo6KFpgF6BJj6GVx748E9cR5OLy5i856R/oDFkDEfsyx6hGW6gL7vGbarxrVsRWi71KA5+9iOGXxTF7I/xPQsQ21dLOQoM6ibQ+7rssPqx58nzyrtdvAu/RXR6CFHxCUOvHQWDg== X-Forefront-Antispam-Report: CIP:216.228.117.160; CTRY:US; LANG:en; SCL:1; SRV:; IPV:NLI; SFV:NSPM; H:mail.nvidia.com; PTR:dc6edge1.nvidia.com; CAT:NONE; SFS:(13230040)(376014)(82310400026)(36860700016)(1800799024)(23010399003)(5023799004)(11063799006)(56012099006)(3023799007)(18002099003)(6133799003)(13003099007); DIR:OUT; SFP:1101; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: mJr7Oj/Y08doI7IctCqg5l/mF2B2C5aPPdfwDIAxlLpS25lnha/I1/sHNlZWOaLkEa/lzCc26yb8nIh85irfI5bJiKnNq9XhXw8f3NTt8lpsIBi7qIhjx3Ocs9v+srPA2Zz2S5PtZM+CM0JUSjP2hzzS6uuNvZM4X0ldFr4CnZro1WT4AzoVDpYm/PsbnR7S3vB5bvfkNHLYcKJSh87SKM30WgLHrjY9IOx9VO92/tJZxFwpxT3tbXfIjzLqB6ejVr/RyjRqspn6Yuj1mzcFg6GLt06ojIyMV+zu0dikhsuNCVxNyl17nu5qK/aFQE+KWBMhwlQIJk2oECmaVstaBs8DSMU2Gz+PLtJNqfKHbS+u/finB3ggjb8VekzXZqeIcHC9haubBimZw4gANx/TAnLbHta1k9GPdST/CIo/2Y83Cel1iq5bVuGrpAnuNm0m X-OriginatorOrg: Nvidia.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 08 Jul 2026 06:05:12.6341 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: b8559cb8-10f7-4a87-bf03-08dedcb6dfa9 X-MS-Exchange-CrossTenant-Id: 43083d15-7273-40c1-b7db-39efd9ccc17a X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=43083d15-7273-40c1-b7db-39efd9ccc17a; Ip=[216.228.117.160]; Helo=[mail.nvidia.com] X-MS-Exchange-CrossTenant-AuthSource: DS1PEPF00017095.namprd03.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: SJ1PR12MB6100 X-Spam-Status: No, score=-8.0 required=5.0 tests=BAYES_00, DKIMWL_WL_HIGH, DKIM_SIGNED, DKIM_VALID, DKIM_VALID_AU, DKIM_VALID_EF, FORGED_SPF_HELO, GIT_PATCH_0, LOCAL_AUTHENTICATION_FAIL_SPF, RCVD_IN_DNSWL_NONE, RCVD_IN_MSPIKE_H2, SPF_HELO_PASS, SPF_NONE, TXREP shortcircuit=no autolearn=ham autolearn_force=no version=3.4.6 X-Spam-Checker-Version: SpamAssassin 3.4.6 (2021-04-09) on sourceware.org X-BeenThere: gcc-patches@gcc.gnu.org X-Mailman-Version: 2.1.30 Precedence: list List-Id: Gcc-patches mailing list List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: gcc-patches-bounces~patchwork=sourceware.org@gcc.gnu.org From: Kyrylo Tkachov A comparison mask m = cmp ? CST1 : CST2 feeding several tests if (m == C) is kept as a constant-pair PHI, so each test is a separate diamond and the backward jump threader tail-duplicates the chain; the selects then come out as branches instead of conditional moves (PR125672). Let phiopt_early_allow accept the CMP, (convert)(compare) and -(convert)(compare) forms: sequences consisting only of conversions, negations and comparisons. An equality test of the materialized mask folds to the comparison itself, exposing the shared condition for the pass_merge_diamonds change that follows. match_simplify_replacement gains two checks for the new forms: * the user predictor check now also covers the multi-statement mask sequences; * a mask is not materialized for a diamond that other paths enter, through its join or through a shared arm. That would leave a PHI mixing the mask with the other paths' values, which nothing downstream recovers, and lose the ccmp chains tree-ifcombine forms from the original branches. Testsuite: pr21090.c, pr68198.c, predict-15.c and pr109071_6.c keep testing their original shapes with phiopt1 disabled, and new pr21090-2.c covers the improved folding; vrp102.c scans evrp, where the tested property is now visible; auto-init-uninit-15.c now warns at the real use instead of the inlined location, so its previously xfailed alternative becomes the expectation. Run-time neutral on 731.astcenc_r by itself; it enables the 20% improvement delivered by the following patch. Bootstrapped and regression tested on aarch64-unknown-linux-gnu. gcc/ChangeLog: PR tree-optimization/125672 PR tree-optimization/126043 * tree-ssa-phiopt.cc (phiopt_early_allow): Accept CMP, (convert)(compare) and -(convert)(compare) sequences consisting only of conversions, negations and comparisons. (match_simplify_replacement): Extend the predictor keeping check to the new mask forms. Reject materializing a mask for a diamond whose join has additional predecessors or whose arm is shared with another path. gcc/testsuite/ChangeLog: PR tree-optimization/125672 PR tree-optimization/126043 * gcc.dg/tree-ssa/pr125672-1.c: New test. * gcc.dg/tree-ssa/pr21090-2.c: New test. * gcc.dg/tree-ssa/pr21090.c: Disable phiopt1. * gcc.dg/tree-ssa/pr68198.c: Likewise. * gcc.dg/predict-15.c: Likewise. * gcc.dg/pr109071_6.c: Likewise. * gcc.dg/tree-ssa/vrp102.c: Scan evrp instead of vrp1. * gcc.dg/auto-init-uninit-15.c: Expect the warning at the use in bar instead of the inlined location. Signed-off-by: Kyrylo Tkachov --- gcc/testsuite/gcc.dg/auto-init-uninit-15.c | 12 ++-- gcc/testsuite/gcc.dg/pr109071_6.c | 5 +- gcc/testsuite/gcc.dg/predict-15.c | 4 +- gcc/testsuite/gcc.dg/tree-ssa/pr125672-1.c | 42 ++++++++++++ gcc/testsuite/gcc.dg/tree-ssa/pr21090-2.c | 27 ++++++++ gcc/testsuite/gcc.dg/tree-ssa/pr21090.c | 5 +- gcc/testsuite/gcc.dg/tree-ssa/pr68198.c | 4 +- gcc/testsuite/gcc.dg/tree-ssa/vrp102.c | 4 +- gcc/tree-ssa-phiopt.cc | 80 ++++++++++++++++++---- 9 files changed, 158 insertions(+), 25 deletions(-) create mode 100644 gcc/testsuite/gcc.dg/tree-ssa/pr125672-1.c create mode 100644 gcc/testsuite/gcc.dg/tree-ssa/pr21090-2.c diff --git a/gcc/testsuite/gcc.dg/auto-init-uninit-15.c b/gcc/testsuite/gcc.dg/auto-init-uninit-15.c index 121f0cff274..8d5fdc233b3 100644 --- a/gcc/testsuite/gcc.dg/auto-init-uninit-15.c +++ b/gcc/testsuite/gcc.dg/auto-init-uninit-15.c @@ -1,16 +1,14 @@ /* PR tree-optimization/17506 - We issue an uninitialized variable warning at a wrong location at - line 11, which is very confusing. Make sure we print out a note to - make it less confusing. (not xfailed alternative) - But it is of course ok if we warn in bar about uninitialized use - of j. (not xfailed alternative) */ + We used to issue an uninitialized variable warning at a wrong (inlined) + location, which was very confusing. Since early phiopt materializes + foo's comparison the warning is issued at the real use in bar, with the note pointing at the declaration. */ /* { dg-do compile } */ /* { dg-options "-O1 -Wuninitialized -ftrivial-auto-var-init=zero" } */ inline int foo (int i) { - if (i) /* { dg-warning "used uninitialized" } */ + if (i) return 1; return 0; } @@ -21,6 +19,6 @@ void bar (void) { int j; /* { dg-message "note: 'j' was declared here" "" } */ - for (; foo (j); ++j) /* { dg-warning "'j' is used uninitialized" "" { xfail *-*-* } } */ + for (; foo (j); ++j) /* { dg-warning "'j' is used uninitialized" } */ baz (); } diff --git a/gcc/testsuite/gcc.dg/pr109071_6.c b/gcc/testsuite/gcc.dg/pr109071_6.c index eddf15b350c..6062882583f 100644 --- a/gcc/testsuite/gcc.dg/pr109071_6.c +++ b/gcc/testsuite/gcc.dg/pr109071_6.c @@ -1,7 +1,10 @@ /* PR tree-optimization/109071 need more context for -Warray-bounds warnings due to code duplication from jump threading. test case is from PR117179, which is a duplication of PR109071. */ -/* { dg-options "-O2 -Warray-bounds -fdiagnostics-show-context=1" } */ +/* { dg-options "-O2 -Warray-bounds -fdiagnostics-show-context=1 -fdisable-tree-phiopt1" } */ +/* phiopt1 is disabled: early phiopt now materializes setval_internal's + comparison, so no threaded path with the out-of-bounds range reaches + the array access and no warning is issued at all. */ /* { dg-additional-options "-fdiagnostics-show-line-numbers -fdiagnostics-path-format=inline-events -fdiagnostics-show-caret" } */ /* { dg-enable-nn-line-numbers "" } */ const char* commands[] = {"a", "b"}; diff --git a/gcc/testsuite/gcc.dg/predict-15.c b/gcc/testsuite/gcc.dg/predict-15.c index 2a8c3ea8597..9a25537c9b6 100644 --- a/gcc/testsuite/gcc.dg/predict-15.c +++ b/gcc/testsuite/gcc.dg/predict-15.c @@ -1,5 +1,7 @@ /* { dg-do compile } */ -/* { dg-options "-O2 -fdump-tree-profile_estimate" } */ +/* { dg-options "-O2 -fdump-tree-profile_estimate -fdisable-tree-phiopt1" } */ +/* phiopt1 is disabled: early phiopt now materializes the comparison + and no goto edge remains to predict. */ int main(int argc, char **argv) { diff --git a/gcc/testsuite/gcc.dg/tree-ssa/pr125672-1.c b/gcc/testsuite/gcc.dg/tree-ssa/pr125672-1.c new file mode 100644 index 00000000000..6cc90f656c5 --- /dev/null +++ b/gcc/testsuite/gcc.dg/tree-ssa/pr125672-1.c @@ -0,0 +1,42 @@ +/* Early phiopt materializes a comparison mask "cmp ? CST1 : CST2" as + -(int)cmp / (int)cmp, exposing its comparison. An equality test of the + mask ("m == C") then folds to the comparison itself, so the mask + disappears and diamonds controlled by the same mask share an identical + condition (a prerequisite for pass_merge_diamonds). */ +/* { dg-do compile } */ +/* { dg-options "-O2 -fdump-tree-phiopt1" } */ + +/* -(int)cmp all-ones mask tested against a constant: the test folds to the + comparison. */ +int +neg_masktest (int a, int b, int x, int y) +{ + int m; + if (a > b) m = -1; else m = 0; + return (m == -1) ? x : y; +} + +/* (int)cmp boolean mask tested against a constant (the convert case). */ +int +cvt_masktest (int a, int b, int x, int y) +{ + int m; + if (a > b) m = 1; else m = 0; + return (m == 1) ? x : y; +} + +/* A mask consumed by other operators is materialized too. */ +int +mask_algebra (int a, int b) +{ + int m; + if (a > b) m = -1; else m = 0; + return m & 7; +} + +/* All three masks are materialized: no constant-pair mask PHI survives and + the equality tests fold away. */ +/* { dg-final { scan-tree-dump-not "= PHI <-1\\(" "phiopt1" } } */ +/* { dg-final { scan-tree-dump-not "== -1" "phiopt1" } } */ +/* { dg-final { scan-tree-dump-not "== 1" "phiopt1" } } */ +/* { dg-final { scan-tree-dump "& 7" "phiopt1" } } */ diff --git a/gcc/testsuite/gcc.dg/tree-ssa/pr21090-2.c b/gcc/testsuite/gcc.dg/tree-ssa/pr21090-2.c new file mode 100644 index 00000000000..d3792d4314f --- /dev/null +++ b/gcc/testsuite/gcc.dg/tree-ssa/pr21090-2.c @@ -0,0 +1,27 @@ +/* With early phiopt lowering "p != 0 ? 1 : 0" to + a boolean mask, the non-null test of a PHI of two object addresses folds + away entirely, without needing VRP's predicate folding (compare + pr21090.c). */ +/* { dg-do compile } */ +/* { dg-options "-O2 -fno-thread-jumps -fdelete-null-pointer-checks -fdump-tree-optimized" } */ + +int g, h; + +int +foo (int a) +{ + int *p; + + if (a) + p = &g; + else + p = &h; + + if (p != 0) + return 1; + else + return 0; +} + +/* { dg-final { scan-tree-dump "return 1;" "optimized" { target { ! keeps_null_pointer_checks } } } } */ +/* { dg-final { scan-tree-dump-not "PHI" "optimized" { target { ! keeps_null_pointer_checks } } } } */ diff --git a/gcc/testsuite/gcc.dg/tree-ssa/pr21090.c b/gcc/testsuite/gcc.dg/tree-ssa/pr21090.c index 92a87688601..ef755e85270 100644 --- a/gcc/testsuite/gcc.dg/tree-ssa/pr21090.c +++ b/gcc/testsuite/gcc.dg/tree-ssa/pr21090.c @@ -1,5 +1,8 @@ /* { dg-do compile } */ -/* { dg-options "-O2 -fno-thread-jumps -fdisable-tree-evrp -fdump-tree-vrp1 -fdelete-null-pointer-checks" } */ +/* { dg-options "-O2 -fno-thread-jumps -fdisable-tree-evrp -fdump-tree-vrp1 -fdelete-null-pointer-checks -fdisable-tree-phiopt1" } */ +/* phiopt1 is disabled so the non-null test reaches VRP as a branch; early + phiopt now lowers it to a boolean mask that is folded even earlier, + which pr21090-2.c covers. */ int g, h; diff --git a/gcc/testsuite/gcc.dg/tree-ssa/pr68198.c b/gcc/testsuite/gcc.dg/tree-ssa/pr68198.c index 125072941da..3b33961268f 100644 --- a/gcc/testsuite/gcc.dg/tree-ssa/pr68198.c +++ b/gcc/testsuite/gcc.dg/tree-ssa/pr68198.c @@ -1,5 +1,7 @@ /* { dg-do compile } */ -/* { dg-options "-O2 -fdump-tree-threadfull1-details -fdisable-tree-ethread" } */ +/* { dg-options "-O2 -fdump-tree-threadfull1-details -fdisable-tree-ethread -fdisable-tree-phiopt1" } */ +/* phiopt1 is disabled: early phiopt now lowers the remove test to a boolean + mask and these threading opportunities change shape. */ extern void abort (void); diff --git a/gcc/testsuite/gcc.dg/tree-ssa/vrp102.c b/gcc/testsuite/gcc.dg/tree-ssa/vrp102.c index fb62e570bed..83f6d98796f 100644 --- a/gcc/testsuite/gcc.dg/tree-ssa/vrp102.c +++ b/gcc/testsuite/gcc.dg/tree-ssa/vrp102.c @@ -1,5 +1,5 @@ /* { dg-do compile } */ -/* { dg-options "-O2 -fno-tree-dominator-opts -fdump-tree-vrp1" } */ +/* { dg-options "-O2 -fno-tree-dominator-opts -fdump-tree-evrp" } */ int f(int x, int y) { @@ -15,4 +15,4 @@ int f(int x, int y) /* We should have computed x ^ y as zero and propagated the result into the PHI feeding the result. */ -/* { dg-final { scan-tree-dump "ret_\[0-9\]+ = PHI <\[01\]\\\(\[0-9\]+\\\), \[01\]\\\(\[0-9\]+\\\)>" "vrp1" } } */ +/* { dg-final { scan-tree-dump "ret_\[0-9\]+ = PHI <\[01\]\\\(\[0-9\]+\\\), \[01\]\\\(\[0-9\]+\\\)>" "evrp" } } */ diff --git a/gcc/tree-ssa-phiopt.cc b/gcc/tree-ssa-phiopt.cc index e12dc7a8b0c..df658b04cfb 100644 --- a/gcc/tree-ssa-phiopt.cc +++ b/gcc/tree-ssa-phiopt.cc @@ -666,6 +666,29 @@ phiopt_early_allow (gimple_seq &seq, gimple_match_op &op) code = gimple_assign_rhs_code (stmt); return code == MIN_EXPR || code == MAX_EXPR; } + /* Accept CMP, (convert)(compare) and -(convert)(compare) when the + generated stmts are only conversions and comparisons, so a mask such as + "cmp ? -1 : 0" / "cmp ? 1 : 0" can be materialized as -(int)cmp / + (int)cmp, exposing its comparison. This is only the structural test; + match_simplify_replacement additionally keeps user predictors alive + and leaves diamonds that other paths enter alone. */ + if (code == NEGATE_EXPR + || CONVERT_EXPR_CODE_P (code) + || TREE_CODE_CLASS (code) == tcc_comparison) + { + for (gimple_stmt_iterator gsi = gsi_start (seq); !gsi_end_p (gsi); + gsi_next (&gsi)) + { + gimple *s = gsi_stmt (gsi); + if (!is_gimple_assign (s)) + return false; + tree_code scode = gimple_assign_rhs_code (s); + if (!CONVERT_EXPR_CODE_P (scode) + && TREE_CODE_CLASS (scode) != tcc_comparison) + return false; + } + return true; + } return false; } @@ -684,7 +707,8 @@ phiopt_early_allow (gimple_seq &seq, gimple_match_op &op) case FIXED_CST: return true; default: - return false; + /* A bare comparison (CMP) is also allowed. */ + return TREE_CODE_CLASS (code) == tcc_comparison; } } @@ -1101,9 +1125,9 @@ match_simplify_replacement (basic_block cond_bb, basic_block middle_bb, } /* For early phiopt, we don't want to lose user generated predictors - if the phiopt is converting `if (a)` into `a` as that might - be jump threaded later on so we want to keep around the - predictors. */ + if the phiopt is converting `if (a)` into `a` (or into the new + (convert)cmp / -(convert)cmp mask forms) as that might be jump + threaded later on so we want to keep around the predictors. */ if (early_p && result && TREE_CODE (result) == SSA_NAME) { bool check_it = false; @@ -1111,15 +1135,33 @@ match_simplify_replacement (basic_block cond_bb, basic_block middle_bb, tree cmp1 = gimple_cond_rhs (stmt); if (result == cmp0 || result == cmp1) check_it = true; - else if (gimple_seq_singleton_p (seq)) - { - gimple *stmt = gimple_seq_first_stmt (seq); - if (is_gimple_assign (stmt) - && result == gimple_assign_lhs (stmt) - && TREE_CODE_CLASS (gimple_assign_rhs_code (stmt)) - == tcc_comparison) + else if (!gimple_seq_empty_p (seq)) + { + /* A comparison mask: the generated statements are only + conversions, negations and comparisons, with at least one + comparison (the forms phiopt_early_allow accepts). */ + bool has_cmp = false, only_mask_ops = true; + for (gimple_stmt_iterator gi = gsi_start (seq); !gsi_end_p (gi); + gsi_next (&gi)) + { + gimple *gs = gsi_stmt (gi); + if (!is_gimple_assign (gs)) + { + only_mask_ops = false; + break; + } + tree_code c = gimple_assign_rhs_code (gs); + if (TREE_CODE_CLASS (c) == tcc_comparison) + has_cmp = true; + else if (!CONVERT_EXPR_CODE_P (c) && c != NEGATE_EXPR) + { + only_mask_ops = false; + break; + } + } + if (only_mask_ops && has_cmp) check_it = true; - } + } if (!check_it) ; else if (contains_hot_cold_predict (middle_bb)) @@ -1128,6 +1170,20 @@ match_simplify_replacement (basic_block cond_bb, basic_block middle_bb, && middle_bb != middle_bb_alt && contains_hot_cold_predict (middle_bb_alt)) return false; + /* Materializing a comparison mask for a diamond that other paths + enter (through its join or through a shared arm) would leave + a PHI mixing the mask with the other paths' values: a residual + half-diamond that neither ifcombine nor RTL if-conversion + recovers, whereas they do handle the original nested branches + (e.g. into ccmp chains). Leave such embedded diamonds alone. */ + else if (result != cmp0 && result != cmp1 + && (EDGE_COUNT (gimple_bb (phi)->preds) != 2 + || (middle_bb != gimple_bb (phi) + && !single_pred_p (middle_bb)) + || (threeway_p + && middle_bb_alt != gimple_bb (phi) + && !single_pred_p (middle_bb_alt)))) + return false; } if (!result) From patchwork Wed Jul 8 06:04:37 2026 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Kyrylo Tkachov X-Patchwork-Id: 138710 Return-Path: X-Original-To: patchwork@sourceware.org Delivered-To: patchwork@sourceware.org Received: from vm01.sourceware.org (localhost [IPv6:::1]) by sourceware.org (Postfix) with ESMTP id 4A1124BA2E0B for ; Wed, 8 Jul 2026 06:06:58 +0000 (GMT) DKIM-Filter: OpenDKIM Filter v2.11.0 sourceware.org 4A1124BA2E0B Authentication-Results: sourceware.org; dkim=pass (2048-bit key, unprotected) header.d=Nvidia.com header.i=@Nvidia.com header.a=rsa-sha256 header.s=selector2 header.b=nnSO/6pQ X-Original-To: gcc-patches@gcc.gnu.org Delivered-To: gcc-patches@gcc.gnu.org Received: from BN8PR05CU002.outbound.protection.outlook.com (mail-eastus2azon11011021.outbound.protection.outlook.com [52.101.57.21]) by sourceware.org (Postfix) with ESMTPS id 1C9D84BA5436 for ; Wed, 8 Jul 2026 06:05:22 +0000 (GMT) DMARC-Filter: OpenDMARC Filter v1.4.2 sourceware.org 1C9D84BA5436 Authentication-Results: sourceware.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com Authentication-Results: sourceware.org; spf=fail smtp.mailfrom=nvidia.com ARC-Filter: OpenARC Filter v1.0.0 sourceware.org 1C9D84BA5436 Authentication-Results: sourceware.org; arc=pass smtp.remote-ip=52.101.57.21 ARC-Seal: i=2; a=rsa-sha256; d=sourceware.org; s=key; t=1783490722; cv=pass; b=fMVczDCmNRBcNwu2UfZQUn/pcTBfilNBS+wA54CPOYweFupB0Elal1SxdNxPN0WS+33wAqLtibuixuV5yrMBaWJOkJ+jOxlTU39nB9BB0U5O3jxJ4Rp1vl1lxspfbYGunkIjzLWPEuv8DebouGHGtI7syFSkP9o1Og3GFfo2YFc= ARC-Message-Signature: i=2; a=rsa-sha256; d=sourceware.org; s=key; t=1783490722; c=relaxed/simple; bh=KR8ZQ9UcQp/UZJaOtny+6+uNpynuKNSAQ6qFOuxCVr8=; h=DKIM-Signature:From:To:Subject:Date:Message-ID:MIME-Version; b=TFbztDcO/vHRCwVIYAcUl9CMJVuYKt9d47LyH/3Kb/777ELySmKQz5WgbdBdc8g89fEslQW152iAh1kDGN8sukfg7UF8CKHKKWABfE89p3T1H1nsQJW38n2csVJqzgxcMw7Cls1SZWnvO6e+YMfx3OGOJqJ2nqlhs86Shg3S5r4= ARC-Authentication-Results: i=2; sourceware.org; dkim=pass (2048-bit key, unprotected) header.d=Nvidia.com header.i=@Nvidia.com header.a=rsa-sha256 header.s=selector2 header.b=nnSO/6pQ DKIM-Filter: OpenDKIM Filter v2.11.0 sourceware.org 1C9D84BA5436 ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=CUP7tKXGHaj4zNiD8TnFgmPoG0mQSYH9MkIlUGFsEs+v1NEgSl5FNPdhmlpHNQxyDyTpBkzPem8MpGsHqD0M+DnsDP1a3CYHxFUSNiwTq2ORehRbaUL217238a9Bnn9Oy+9NC66WhPtIOnv8zbdsgYJ97YWNmws8cl2lku6z01ljGWV8OWlKjzZiFlYFxvzNsJBYlxxJxCGj78+5LwZTa7YrouU7U9yMIfXJLApn8SXjNZI8b1Vq9OVxatdVuBJTukDXgoPG91pUgBDIYNGppk1Q8x0WSMr8QRARDh4BXgikLjr9D8NThOaOjdPFnbpGQyxPdq1DyobBiRt2iGsXDw== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=paWI7ChIvXuC62fmdAApK4fEvBtN0FYlwLp53N+BnTM=; b=IVdXIY+PfNu15MeeUNtLJwZikPSm6DBPRdQmi/wBpL83aj4xsOgyw+9CCedyUS2BLlrDP5yFqhnBiQUw5efCPEXo49qYTznDGlYmOBLPYcmY4iYJvfXeN9xTRGezBSWEaZjFHFkJn9f1jMxp+dom9UGoQQPlfXw/WsAc3Md72PmOVmX1jzAaBEBq6ijkG9/krx+rFseeWgcQNpN+K1uDEpyGri77P3/xk1smRFd4m2z9l4Lg/RWAQgVqSSuOdCvtWhxtLK+dTxEuSmiB7jKFY4DcCFdLnrHmcMIiuSA214lLY/mQ9xarp0t2Bl8i4LmWo5wodlKXEu4KjQbHviLVMA== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 216.228.117.160) smtp.rcpttodomain=gcc.gnu.org smtp.mailfrom=nvidia.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=nvidia.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=Nvidia.com; s=selector2; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=paWI7ChIvXuC62fmdAApK4fEvBtN0FYlwLp53N+BnTM=; b=nnSO/6pQoS7hXjeNouX//Yf9cpN7CnladGnWkavxnYG4PLscd5sphlaVhhFfmpm7tmF+ODWGhJ/Oms8M7AAgP8U+/C/66/bgjz4CRmNedIQV/0GHCi3zVSEypWkwoibS+kZ6bI3PGsNh3dfFpMiFe/9pGgbajEPjtq/giCgahGhR+/NHM3AKZA9zEP/OTtvo8mzyRnEPIEpHBxPdNztF7sREmFfgyruevLJHYqjZxNC8PlmFbqXCXXu1A7ZUmjhn6ktSLrEik/OE8cyBgEuPS3Lq1YKh3YEOFrneSoI8QKwQTxKmP7BDlH3AHgaRK6p8WLzfW+C/bKYop4/oh8zhXw== Received: from DM6PR02CA0136.namprd02.prod.outlook.com (2603:10b6:5:1b4::38) by BN3PR12MB9572.namprd12.prod.outlook.com (2603:10b6:408:2ca::17) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.181.11; Wed, 8 Jul 2026 06:05:15 +0000 Received: from DS1PEPF0001708E.namprd03.prod.outlook.com (2603:10b6:5:1b4:cafe::a4) by DM6PR02CA0136.outlook.office365.com (2603:10b6:5:1b4::38) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.181.13 via Frontend Transport; Wed, 8 Jul 2026 06:05:15 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 216.228.117.160) smtp.mailfrom=nvidia.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=nvidia.com; Received-SPF: Pass (protection.outlook.com: domain of nvidia.com designates 216.228.117.160 as permitted sender) receiver=protection.outlook.com; client-ip=216.228.117.160; helo=mail.nvidia.com; pr=C Received: from mail.nvidia.com (216.228.117.160) by DS1PEPF0001708E.mail.protection.outlook.com (10.167.17.134) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.181.6 via Frontend Transport; Wed, 8 Jul 2026 06:05:15 +0000 Received: from rnnvmail201.nvidia.com (10.129.68.8) by mail.nvidia.com (10.129.200.66) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.20; Tue, 7 Jul 2026 23:04:57 -0700 Received: from ktkachov-mlt.nvidia.com (10.126.231.37) by rnnvmail201.nvidia.com (10.129.68.8) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.20; Tue, 7 Jul 2026 23:04:55 -0700 From: To: CC: , , Kyrylo Tkachov Subject: [PATCH 2/2] tree-optimization/125672 - merge same-condition if-convertible diamonds Date: Wed, 8 Jul 2026 08:04:37 +0200 Message-ID: <20260708060437.26491-2-ktkachov@nvidia.com> X-Mailer: git-send-email 2.50.1 In-Reply-To: <20260708060437.26491-1-ktkachov@nvidia.com> References: <20260708060437.26491-1-ktkachov@nvidia.com> MIME-Version: 1.0 X-Originating-IP: [10.126.231.37] X-ClientProxiedBy: rnnvmail202.nvidia.com (10.129.68.7) To rnnvmail201.nvidia.com (10.129.68.8) X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: DS1PEPF0001708E:EE_|BN3PR12MB9572:EE_ X-MS-Office365-Filtering-Correlation-Id: 0cc4e15a-dd29-409c-575c-08dedcb6e11d X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0; ARA:13230040|23010399003|376014|36860700016|1800799024|82310400026|13003099007|6133799003|56012099006|11063799006|5023799004|22082099003|18002099003|3023799007; X-Microsoft-Antispam-Message-Info: qeSHI6QQe0KyxhXiJ4TJU9tQl4bNFLU+31/u/E6aKm1BPWG5ZK43Q7JGujxv9P2XADJv83e4ImLsI3SQsnLPS61Bw/goxAJ65n9KDW9guUhx8H/Lk79Oud8ahQBGz8kvp/V57gwIEWQ+ziyLrECA2K79Xd6py8VSHaLhjfRph+P/CACzvImdaPxZ9A/WpeSrOBCbI3XcxgcRyNPLjL4uiw6pkgLRmlYHZbbr+xthxmxAs8Xt5qUAUcN67kodaFsiRBQ3+5cZdOpo8yrPgsLf1KMoGimofp1jGr2wnJ+Dw4OCOyvdFXfUYP/6aYl+qj9V35lVE3LTivEPgU7B9zEovSpLKIoj7j/B4eNtKLlfEuZi8JQCat6j/+A762avMHmwHTp1Bdo8hAm9DspKIEz8EcwBCcnC0M6B9xBFYUe5gT0DzYQhKJAo8QjqagWea08WKK3yjoWt8ZVz3htviHvIa/bu/QU41J8CKH6xtmvOcshXDB/S3dUzWvmNd0AuGzxdvvMpVEPDOzigMylRUJsPWg2MX/lmDa8m7+yfEjMQBFx46aX52FhdyPlqObLfXZkwZrsy5ejw4lR2XtGPIfSkixKG5Rp8Yhs7LsVJ3P/q+TThMM5zD333BmC68jQLnRdufIf73koTvZ3HKq3IvUBADhmZVFC2KJcBgXxFuitjvesk5DEAxdMQONNRlCgd06I/ytJX8MMNVePp485xUFHalQ== X-Forefront-Antispam-Report: CIP:216.228.117.160; CTRY:US; LANG:en; SCL:1; SRV:; IPV:NLI; SFV:NSPM; H:mail.nvidia.com; PTR:dc6edge1.nvidia.com; CAT:NONE; SFS:(13230040)(23010399003)(376014)(36860700016)(1800799024)(82310400026)(13003099007)(6133799003)(56012099006)(11063799006)(5023799004)(22082099003)(18002099003)(3023799007); DIR:OUT; SFP:1101; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: 0j67ggrXIJl8249F53Q2/aZqB9RU8a5s+Ij5UXHB+Vs2OCA93wPRbb1E7p5BYdmKhGMkjZ+MI9zAkHKQs6i0g0D3S5EKfF3+J7wXsMRyipkc4lRCwyKgQuqufe4cUqx6fII5M3RXBXvxv6HeAsSubPxqI2TKaCQGuMayooTnEeexiNDist4xD41nKMP1E8M/XSNzfgs8jC9lJjglwxK2/jYjSqt6qAEjRKGUb6ygaLqCO4iArx7HBf1AkLmf5RjCYOYQYPobU+tf8D/w/KDg//YFMIlgGJOqrVtq1Ho6gtE15LThK4bpmuJtqed1sDrkCsR4oP4FjUI+P1AR8NYOB4ZZDSnQSn9JJ/t6xR/pl941EwVhe/Zzl2DdBnsqj2g7Qakg3ct3uB7kHGmm82b00pHUoHfY+too77JS4cvexLmBQMFaOKXANCSi++z8Mrgr X-OriginatorOrg: Nvidia.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 08 Jul 2026 06:05:15.0508 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: 0cc4e15a-dd29-409c-575c-08dedcb6e11d X-MS-Exchange-CrossTenant-Id: 43083d15-7273-40c1-b7db-39efd9ccc17a X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=43083d15-7273-40c1-b7db-39efd9ccc17a; Ip=[216.228.117.160]; Helo=[mail.nvidia.com] X-MS-Exchange-CrossTenant-AuthSource: DS1PEPF0001708E.namprd03.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: BN3PR12MB9572 X-Spam-Status: No, score=-8.0 required=5.0 tests=BAYES_00, DKIMWL_WL_HIGH, DKIM_SIGNED, DKIM_VALID, DKIM_VALID_AU, DKIM_VALID_EF, FORGED_SPF_HELO, GIT_PATCH_0, LOCAL_AUTHENTICATION_FAIL_SPF, RCVD_IN_DNSWL_NONE, RCVD_IN_MSPIKE_H2, SPF_HELO_PASS, SPF_NONE, TXREP shortcircuit=no autolearn=ham autolearn_force=no version=3.4.6 X-Spam-Checker-Version: SpamAssassin 3.4.6 (2021-04-09) on sourceware.org X-BeenThere: gcc-patches@gcc.gnu.org X-Mailman-Version: 2.1.30 Precedence: list List-Id: Gcc-patches mailing list List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: gcc-patches-bounces~patchwork=sourceware.org@gcc.gnu.org From: Kyrylo Tkachov When a chain of if-convertible diamonds shares a controlling condition, the backward jump threader tail-duplicates it into a 2^N decision tree and the per-lane selects can no longer be if-converted. Add pass_merge_diamonds, run before pass_thread_jumps_full and gated on flag_thread_jumps. For each if-convertible diamond it walks the immediate dominators for an earlier diamond with an identical GIMPLE_COND and, when all non-virtual join PHIs of the later diamond have their arguments available at the earlier join, moves the PHIs up into it, keeping their result SSA names, and removes the now redundant branch. The threader then finds no correlated conditions to duplicate and RTL if-conversion lowers the selects to conditional moves; whether each select becomes a move remains a cost-model decision. The scan collects candidate blocks by index up front and processes them in one sweep, looking each up again by index since a merge deletes blocks. A chain of same-condition diamonds merges bottom-up within a single sweep (remove_edge_and_dominated_blocks keeps the dominator information valid throughout); a rescan happens only when a merge changed the topology in a way that could enable a merge that failed earlier, so the common case is one sweep plus one empty verification sweep. Improves 731.astcenc_r (SPEC CPU2026) by 20% on aarch64 at -Ofast -flto; the hot select kernel goes from 36 conditional branches and 12 conditional moves to 8 and 24. Placing the pass after the first VRP instance instead gives identical results, so the minimal placement is kept. Bootstrapped and regression tested on aarch64-unknown-linux-gnu. gcc/ChangeLog: PR tree-optimization/125672 * passes.def: Add pass_merge_diamonds before pass_thread_jumps_full. * tree-pass.h (make_pass_merge_diamonds): Declare. * tree-ssa-ifcombine.cc (ifcvt_diamond_join, value_available_at) (abnormal_ssa_p, merge_cond_diamond_up): New functions. (class pass_merge_diamonds): New pass. (make_pass_merge_diamonds): New function. gcc/testsuite/ChangeLog: PR tree-optimization/125672 * gcc.dg/tree-ssa/pr125672-2.c: New test. * gcc.dg/tree-ssa/pr125672-3.c: New test. Signed-off-by: Kyrylo Tkachov --- gcc/passes.def | 1 + gcc/testsuite/gcc.dg/tree-ssa/pr125672-2.c | 17 ++ gcc/testsuite/gcc.dg/tree-ssa/pr125672-3.c | 32 +++ gcc/tree-pass.h | 1 + gcc/tree-ssa-ifcombine.cc | 284 +++++++++++++++++++++ 5 files changed, 335 insertions(+) create mode 100644 gcc/testsuite/gcc.dg/tree-ssa/pr125672-2.c create mode 100644 gcc/testsuite/gcc.dg/tree-ssa/pr125672-3.c diff --git a/gcc/passes.def b/gcc/passes.def index 1fc867fae51..3c4f0629e78 100644 --- a/gcc/passes.def +++ b/gcc/passes.def @@ -233,6 +233,7 @@ along with GCC; see the file COPYING3. If not see NEXT_PASS (pass_return_slot); NEXT_PASS (pass_fre, true /* may_iterate */); NEXT_PASS (pass_merge_phi); + NEXT_PASS (pass_merge_diamonds); NEXT_PASS (pass_thread_jumps_full, /*first=*/true); NEXT_PASS (pass_vrp, false /* final_p*/); NEXT_PASS (pass_array_bounds); diff --git a/gcc/testsuite/gcc.dg/tree-ssa/pr125672-2.c b/gcc/testsuite/gcc.dg/tree-ssa/pr125672-2.c new file mode 100644 index 00000000000..e66eeb0eac1 --- /dev/null +++ b/gcc/testsuite/gcc.dg/tree-ssa/pr125672-2.c @@ -0,0 +1,17 @@ +/* Same-condition if-convertible diamonds separated by intervening diamonds are + merged before jump threading so the threader cannot tail-duplicate them. */ +/* { dg-do compile } */ +/* { dg-options "-O2 -fdump-tree-mergediam-details" } */ + +void f (float *o, float a1, float a2, float b0, float b1, + float c0, float c1, float d0, float d1) +{ + float t0, t1, t2, t3; + if (a1 > 0) t0 = b0; else t0 = c0; + if (a2 > 0) t1 = b1; else t1 = c1; + if (a1 > 0) t2 = d0; else t2 = d1; + if (a2 > 0) t3 = b0; else t3 = c1; + o[0] = t0; o[1] = t1; o[2] = t2; o[3] = t3; +} + +/* { dg-final { scan-tree-dump-times "merging if-convertible diamond" 2 "mergediam" } } */ diff --git a/gcc/testsuite/gcc.dg/tree-ssa/pr125672-3.c b/gcc/testsuite/gcc.dg/tree-ssa/pr125672-3.c new file mode 100644 index 00000000000..24a4f7ce674 --- /dev/null +++ b/gcc/testsuite/gcc.dg/tree-ssa/pr125672-3.c @@ -0,0 +1,32 @@ +/* Execution test: merging same-condition if-convertible diamonds must + preserve semantics. */ +/* { dg-do run } */ +/* { dg-options "-O2" } */ + +__attribute__((noipa)) static void +kern (int n, const int *a, const int *b, int *o) +{ + for (int i = 0; i < n; i++) + { + int m0 = (a[i] > b[i]) ? -1 : 0; /* a comparison mask */ + int m1 = (a[i] < b[i]) ? -1 : 0; /* a second mask */ + o[4*i + 0] = (m0 == -1) ? a[i] : b[i]; /* select under mask 0 */ + o[4*i + 1] = (m1 == -1) ? a[i] : b[i]; /* select under mask 1 */ + o[4*i + 2] = (m0 == -1) ? b[i] : a[i]; /* mask 0 again */ + o[4*i + 3] = (m1 == -1) ? b[i] : a[i]; /* mask 1 again */ + } +} + +int +main (void) +{ + int a[3] = { 5, 2, 9 }; + int b[3] = { 3, 7, 9 }; + int o[12]; + int exp[12] = { 5, 3, 3, 5, 7, 2, 2, 7, 9, 9, 9, 9 }; + kern (3, a, b, o); + for (int i = 0; i < 12; i++) + if (o[i] != exp[i]) + __builtin_abort (); + return 0; +} diff --git a/gcc/tree-pass.h b/gcc/tree-pass.h index b3c97658a8f..1625c1766e0 100644 --- a/gcc/tree-pass.h +++ b/gcc/tree-pass.h @@ -464,6 +464,7 @@ extern gimple_opt_pass *make_pass_phiopt (gcc::context *ctxt); extern gimple_opt_pass *make_pass_forwprop (gcc::context *ctxt); extern gimple_opt_pass *make_pass_phiprop (gcc::context *ctxt); extern gimple_opt_pass *make_pass_tree_ifcombine (gcc::context *ctxt); +extern gimple_opt_pass *make_pass_merge_diamonds (gcc::context *ctxt); extern gimple_opt_pass *make_pass_dse (gcc::context *ctxt); extern gimple_opt_pass *make_pass_nrv (gcc::context *ctxt); extern gimple_opt_pass *make_pass_rename_ssa_copies (gcc::context *ctxt); diff --git a/gcc/tree-ssa-ifcombine.cc b/gcc/tree-ssa-ifcombine.cc index 0c3e78ef331..fc66ccac81e 100644 --- a/gcc/tree-ssa-ifcombine.cc +++ b/gcc/tree-ssa-ifcombine.cc @@ -1450,3 +1450,287 @@ make_pass_tree_ifcombine (gcc::context *ctxt) { return new pass_tree_ifcombine (ctxt); } + +/* Merge if-convertible diamonds that share a controlling condition. The + backward jump threader tail-duplicates such a chain into a 2^N decision + tree, so the per-lane selects come out as branches instead of conditional + moves. This pass merges a later diamond into an earlier same-condition + one before threading runs: + + if (c) if (c) + / \ / \ + . . (empty arms) . . + \ / \ / + x = PHI x = PHI + | y = PHI + ... ==> | + | ... + if (c) | + / \ ... + . . + \ / + y = PHI + | + ... + + Both p and q are defined above the first join, so the later join's PHI + can be evaluated there, keyed by the shared condition (p on the true + side, q on the false side); the second branch then decides nothing and + is removed. The threader sees one diamond per condition and finds + nothing to tail-duplicate. This relies on early phiopt having + de-indirected constant-pair masks (cmp ? -1 : 0 -> -(int)cmp) so the + repeated mask tests share an identical GIMPLE_COND. Gated on + flag_thread_jumps. */ + +/* If COND_BB heads an if-then-else with empty (side-effect-free) arms that + reconverge at a single join, return the join and set *E_TRUE / *E_FALSE to + the edges entering it on the true / false side (possibly via one empty + forwarder). Otherwise NULL. */ + +static basic_block +ifcvt_diamond_join (basic_block cond_bb, edge *e_true, edge *e_false) +{ + basic_block tb = NULL, fb = NULL; + if (!recognize_if_then_else (cond_bb, &tb, &fb)) + return NULL; + edge te = find_edge (cond_bb, tb); + edge fe = find_edge (cond_bb, fb); + + edge tj_e = te, fj_e = fe; + basic_block tj = tb, fj = fb; + if (single_pred_p (tb) && single_succ_p (tb) && empty_block_p (tb)) + { + tj = single_succ (tb); + tj_e = single_succ_edge (tb); + } + if (single_pred_p (fb) && single_succ_p (fb) && empty_block_p (fb)) + { + fj = single_succ (fb); + fj_e = single_succ_edge (fb); + } + + basic_block join; + if (tb == fj && tb != cond_bb) + { + join = tb; + *e_true = te; + *e_false = fj_e; + } + else if (fb == tj && fb != cond_bb) + { + join = fb; + *e_true = tj_e; + *e_false = fe; + } + else if (tj == fj && tj != cond_bb && tb != fb) + { + join = tj; + *e_true = tj_e; + *e_false = fj_e; + } + else + return NULL; + + if (EDGE_COUNT (join->preds) != 2 + || (*e_true)->src == (*e_false)->src) + return NULL; + if (tb != join && !bb_no_side_effects_p (tb)) + return NULL; + if (fb != join && !bb_no_side_effects_p (fb)) + return NULL; + return join; +} + +/* True if VAL (a constant, default def, or SSA name whose definition dominates + BB) is available at BB. */ + +static bool +value_available_at (tree val, basic_block bb) +{ + if (TREE_CODE (val) != SSA_NAME || SSA_NAME_IS_DEFAULT_DEF (val)) + return true; + basic_block def_bb = gimple_bb (SSA_NAME_DEF_STMT (val)); + return def_bb && dominated_by_p (CDI_DOMINATORS, bb, def_bb); +} + +/* True if T is an SSA name that occurs in an abnormal PHI (must not be moved + or have its definition relocated). */ + +static inline bool +abnormal_ssa_p (tree t) +{ + return TREE_CODE (t) == SSA_NAME && SSA_NAME_OCCURS_IN_ABNORMAL_PHI (t); +} + +/* Try to merge the same-condition if-convertible diamond headed by B2 into an + earlier dominating one by moving B2's join PHIs up. Return true if done. */ + +static bool +merge_cond_diamond_up (basic_block b2) +{ + edge t2, f2; + basic_block join2 = ifcvt_diamond_join (b2, &t2, &f2); + if (!join2) + return false; + gcond *c2 = safe_dyn_cast (*gsi_last_bb (b2)); + if (!c2) + return false; + tree cl = gimple_cond_lhs (c2), cr = gimple_cond_rhs (c2); + if (abnormal_ssa_p (cl) || abnormal_ssa_p (cr)) + return false; + + for (basic_block b1 = get_immediate_dominator (CDI_DOMINATORS, b2); + b1; b1 = get_immediate_dominator (CDI_DOMINATORS, b1)) + { + gcond *c1 = safe_dyn_cast (*gsi_last_bb (b1)); + if (!c1 + || gimple_cond_code (c1) != gimple_cond_code (c2) + || !operand_equal_p (gimple_cond_lhs (c1), cl, 0) + || !operand_equal_p (gimple_cond_rhs (c1), cr, 0)) + continue; + + edge t1, f1; + basic_block join1 = ifcvt_diamond_join (b1, &t1, &f1); + if (!join1 || join1 == join2 || join1 == b2 || join1 == b1) + continue; + + /* Each non-virtual JOIN2 PHI must have both args available at B1 and be + free of abnormal coalescing, so it can be moved up to JOIN1. */ + bool ok = true, any = false; + for (gphi_iterator gpi = gsi_start_phis (join2); + !gsi_end_p (gpi); gsi_next (&gpi)) + { + gphi *phi = gpi.phi (); + tree res = gimple_phi_result (phi); + if (virtual_operand_p (res)) + continue; + any = true; + tree tv = PHI_ARG_DEF_FROM_EDGE (phi, t2); + tree fv = PHI_ARG_DEF_FROM_EDGE (phi, f2); + if (abnormal_ssa_p (res) || abnormal_ssa_p (tv) || abnormal_ssa_p (fv) + || !value_available_at (tv, b1) || !value_available_at (fv, b1)) + { + ok = false; + break; + } + } + if (!ok || !any) + continue; + + if (dump_file && (dump_flags & TDF_DETAILS)) + fprintf (dump_file, + "merging if-convertible diamond bb%d up into same-condition " + "diamond bb%d\n", b2->index, b1->index); + + /* Move each non-virtual JOIN2 PHI to JOIN1, keeping its result SSA name. + JOIN1 dominates all of the result's uses, so no use or debug bind + changes. The virtual PHI is left to degenerate after cfg cleanup. */ + for (gphi_iterator gpi = gsi_start_phis (join2); !gsi_end_p (gpi);) + { + gphi *phi = gpi.phi (); + tree res = gimple_phi_result (phi); + if (virtual_operand_p (res)) + { + gsi_next (&gpi); + continue; + } + tree tv = PHI_ARG_DEF_FROM_EDGE (phi, t2); + tree fv = PHI_ARG_DEF_FROM_EDGE (phi, f2); + location_t tl = gimple_phi_arg_location_from_edge (phi, t2); + location_t fl = gimple_phi_arg_location_from_edge (phi, f2); + remove_phi_node (&gpi, false); + gphi *nphi = create_phi_node (res, join1); + add_phi_arg (nphi, tv, t1, tl); + add_phi_arg (nphi, fv, f1, fl); + } + + /* B2's branch is now redundant; keep its true edge as a fallthrough and + drop the false edge with any unreachable blocks. */ + edge b2t, b2f; + extract_true_false_edges_from_block (b2, &b2t, &b2f); + gimple_stmt_iterator gsic2 = gsi_last_bb (b2); + gsi_remove (&gsic2, true); + remove_edge_and_dominated_blocks (b2f); + b2t->flags &= ~(EDGE_TRUE_VALUE | EDGE_FALSE_VALUE); + b2t->flags |= EDGE_FALLTHRU; + b2t->probability = profile_probability::always (); + loops_state_set (LOOPS_NEED_FIXUP); + return true; + } + return false; +} + +namespace { + +const pass_data pass_data_merge_diamonds = +{ + GIMPLE_PASS, /* type */ + "mergediam", /* name */ + OPTGROUP_NONE, /* optinfo_flags */ + TV_TREE_IFCOMBINE, /* tv_id */ + PROP_cfg | PROP_ssa, /* properties_required */ + 0, /* properties_provided */ + 0, /* properties_destroyed */ + 0, /* todo_flags_start */ + 0, /* todo_flags_finish */ +}; + +class pass_merge_diamonds : public gimple_opt_pass +{ +public: + pass_merge_diamonds (gcc::context *ctxt) + : gimple_opt_pass (pass_data_merge_diamonds, ctxt) + {} + + bool gate (function *) final override + { + return optimize > 0 && flag_thread_jumps && !optimize_debug; + } + + unsigned int execute (function *fun) final override + { + calculate_dominance_info (CDI_DOMINATORS); + bool changed = false; + /* Collect candidates first and look them up again by index: a merge + deletes basic blocks (the redundant branch, and blocks that become + unreachable), so the block list must not be iterated while + mutating. remove_edge_and_dominated_blocks keeps the dominator + information valid throughout, so all merges can be done in one + sweep; in particular a chain of same-condition diamonds merges + bottom-up in a single pass, each merge carrying its join PHIs to + the next. A merge can also change a join's predecessor count and + thereby enable a merge that failed earlier in the sweep, so rescan + until a sweep performs no merge; almost always the second sweep + finds nothing. */ + auto_vec candidates; + bool again = true; + while (again) + { + again = false; + candidates.truncate (0); + basic_block bb; + FOR_EACH_BB_FN (bb, fun) + if (safe_is_a (*gsi_last_bb (bb))) + candidates.safe_push (bb->index); + unsigned i; + int idx; + FOR_EACH_VEC_ELT (candidates, i, idx) + { + basic_block b2 = BASIC_BLOCK_FOR_FN (fun, idx); + if (!b2 || !safe_is_a (*gsi_last_bb (b2))) + continue; + if (merge_cond_diamond_up (b2)) + changed = again = true; + } + } + return changed ? TODO_cleanup_cfg : 0; + } +}; // class pass_merge_diamonds + +} // anon namespace + +gimple_opt_pass * +make_pass_merge_diamonds (gcc::context *ctxt) +{ + return new pass_merge_diamonds (ctxt); +}