Skip to content

ENH: Investigate NPY_GCC_OPT_3 macro for other compilers #21164

Description

@seberg

Since a very long time (predating clang, I believe), NumPy uses NPY_GCC_OPT_3 in a few places. This macro is useful to locally enable a high optimization level for functions we know should be optimized: tight, simple (usually 1-D) loops.

However, due to its age, the macro only applies to GCC (unless clang picks it up?). It would be nice to generalize the macro a bit to other compilers, probably using #pragma depending on the compilers.
This may need some care, since different compilers have different ideas of what O3 means, IIRC. So it may be that e.g. the Intel compiler enables unsafe fast-math when GCC does not.

The task are:

  1. Check how various compilers change the optimization level for a single function
  2. Add additional branches for those compilers to the #define (maybe renaming it)
  3. Check benchmarks of functions that should modified it with the compilers in question
  4. Run the test suite
  5. Double check the compiler documentation to be sure that no unsafe fast-math is enabled. We cannot trust our test-suite on all accounts (e.g. floating point error flags).

In some cases functions that currently use this, may end up as universal-intrinsics eventually. But I somewhat expect that this macro will stay useful for things where maximum performance is less important or just as a stop gap, because it adds no complexity.

EDIT: I expect this is a fairly nice sprintable thing to investigate, although it is best if an MSVC setup is available. (Clang likely supports the gcc attributes, but I am not sure. This is a useful reference probably: https://stackoverflow.com/questions/31373885/how-to-change-optimization-level-of-one-function.)

Activity

  1. added
    sprintableIssue fits the time-frame and setting of a sprint
    on Apr 28, 2022
  2. mar-galho commented on Jun 22, 2022

    @mar-galho

    Hi Sebastian,

    I've started digging into it. Can we hop in a meeting someday? I have a few questions related to this issue.

  3. mar-galho commented on Jul 22, 2023

    @mar-galho

    @seberg
    A year later, but not forgotten. Now, with more time on my hands!

    Is something like this that we'd be looking for?

    #ifdef _MSC_VER
    #define NPY_OPT_3 __pragma(optimize("gt", on))
    #elif defined(__GNUC__) || defined(__clang__)
    #define NPY_OPT_3 __attribute__((optimize("O3")))
    #else
    #define NPY_OPT_3
    #endif

    I can set up a few environments (Docker, possibly, because of personal laziness...even though there's the overhead associated with it that would have to be considered) to benchmark MSVC, Clang, GCC, and ICC?

    Let me know if that's similar to what you're looking for.

  4. seberg commented on Jul 23, 2023

    @seberg
    MemberAuthor

    Yes, probably that is all there is to it. I think we now compile more parts with high optimization anyway, but there are places that should benefit.

  5. mar-galho commented on Jul 23, 2023

    @mar-galho

    I'll run some benchmarking mumbo jumbo here and post the results to see if the gain is worthwhile.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions