Repository navigation
Conversation
|
Thanks for your pull request! It looks like this may be your first contribution to a Google open source project. Before we can look at your pull request, you'll need to sign a Contributor License Agreement (CLA). View this failed invocation of the CLA check for more information. For the most up to date status, view the checks section at the bottom of the pull request. |
167e123 to
24d50c2
Compare
grebe
left a comment
There was a problem hiding this comment.
First of all: thank you! This is exciting to see. Looks like a lot of good work and a big improvement.
Here's some initial feedback, I'm still working my way through. I think I have enough here to get the ball rolling, though, so I'm sending it out. Feel free to reach out directly or go back and forth here if anything I said was unclear or if you disagree.
There was a problem hiding this comment.
Deps shouldn't be under xls/, they should be managed from dependency_support/ or (ideally) BCR. tinyxml2 looks like it's in BCR. scipy_lsap.cpp is a bit trickier b/c our python deps aren't super well integrated into bazel (it's essentially pip installed, which afaik largely hides the internals). We typically try to avoid holding on to third_party/ source in here. I believe OR-tools has some implementations of linear assignment stuff and we already pull it in as a dep - is a pain to move to using it?
There was a problem hiding this comment.
I initially used OR-tools LASP solver but the performance took a hit. I'll do more experiments and will try to replace the external LSAP solver with something more convenient.
Regarding TinyXML, I think it would be nice to get rid custom parses (ir2nx, ir2gxl) and solely use XLS IR and generate graph directly from it before we merge this PR and after the current code is stable.
There was a problem hiding this comment.
Re: using XLS IR, I think it's fine to land this w/ tinyxml as long as you use the BCR version, but agree it would be better not to have custom parsing.
|
@grebe Ready for next round of review. I addressed most of the comments but the cost functions are still in a separate header, if you prefer them to live inside ged header, I can move them in a follow-up commit. |
There was a problem hiding this comment.
Re: using XLS IR, I think it's fine to land this w/ tinyxml as long as you use the BCR version, but agree it would be better not to have custom parsing.
grebe
left a comment
There was a problem hiding this comment.
Sorry for the long wait.
Hopefully this isn't too annoying, but... I think it would be good to move this into contrib (xls/contrib/eco).
We're getting close!
|
@grebe Sorry for late response I was busy with some academic tasks. I applied most of the recent comments, currently working on moving the P.S. The test |
There was a problem hiding this comment.
This should use BCR rather than checking this in
|
@grebe Do you have a preference here:
Happy to go either way, |
I don't think xls_ir_to_cytoscape obviates the entire set of python tools, just the parser. Right? I weakly prefer to move to xls_ir_to_cytoscape, but don't think it needs to be a huge priority. Using TinyXML is fine, we just generally don't check in third-party code as it's easier to manage through bzlmod/BCR. It's possible you won't even need to update the include paths when switching to BCR. |
actually the whole py part is now implemented in cpp, although getting rid of FYI:
|
|
@grebe I went ahead and added direct XLS IR -> ECO graph conversion for the C++ GED flow. The graph now carries typed XLS attributes ( I can deprecate and remove the Python side completely in the next commit. If complete removal is not preferred, please let me know how you’d like to keep it. |
a369e65 to
ceed0d3
Compare
|
@proppy could you please trigger the CI tests? Thanks! |
grebe
left a comment
There was a problem hiding this comment.
Thanks, and sorry for the slow response. Yeah, if it's not too much work to remove the python, feel free.
|
Some other feedback after chatting with @proppy:
|
|
Sorry for the long delay, I was bogged down with academic tasks. Python side is now removed and everything is cpp. I also added a no, didn't really look at or-tools for this. XLSGraph isn't really a generic graph. Each node carries cost_attributes (op, type, literal, state info) plus a hash label that both the MCS candidate partitioning and the GED cost function read directly. Edges carry operand-position indices because Sub(a, b) ≠ Sub(b, a). And the MCS→GED handoff mutates the graph; boundary pairs get pinned, interior matches get cut, with an original↔current index map kept around. or-tools' fast graphs are CSR-style, immutable, and indexed by (source, target) only, so we'd be carrying all the payload in side-arrays and reimplementing the mutation API. At that point the wrapper is the graph. summary: or-tools is not MCS or GED aware and I wanted minimal and controllable overhead so I opted in with a custom graph data structure. |
- Replace requirement() with @xls_pip_deps// for old Python tools - Update Receive node construction to match new upstream API signature - Add payload_type parameter required by new Receive constructor
This commit contains two types of changes: 1. Upstream updates to ir2nx.py merged during rebase (large diff) 2. Our deprecation TODOs added to Python GED toolchain: - ir_diff.py: NetworkX GED → ged.h/cc (10-100x faster) - ir_diff_main.py: Python tool → ged_main.cc CLI - ir_patch_gen.py: Python patcher → patch_ir.h/cc - ir2nx.py: NetworkX converter → xls_ir_to_cytoscape.cc - ir2gxl.py: GXL exporter → gxl_parser.h/cc (from upstream)
Replace the SciPy LSAP path with a native C++ LSAP solver. Switch GED cost construction from sparse-style handling to dense matrix paths. Add C++ ECO components for GED, MCS, graph modeling, GXL parsing, and patch generation. Introduce eco-specific Starlark build defs and wire patching plus equivalence validation targets. Standardize main-flag handling and add optional execution statistics reporting. Add ECO regression fixtures and test targets for crc32, riscv, apfloat, fir, histogram, and vector_core flows.
Use path-based equivalence report flag and write reports for mismatch results Replace run_shell wrapper with direct Bazel run Remove short flag aliases from patch_ir_main Normalize ECO build rule args/report handling and clean stale target references
Fixed init parsing bug in IrParser Fixed a critical bug in MakePruneFunction
Relocate the ECO package from xls/eco to xls/contrib/eco and update references. This includes Bazel labels, Starlark loads, C++ include paths, Python imports, runfiles paths, and tooling path filters.
- Updated `ir_diff_main.py` to reflect direct XLS IR parsing. - Extended `ir_patch.proto` with new fields: message, label, format, and verbosity. - Modified `ir_patch_gen.cc` to handle new attributes for assert and trace operations. - Enhanced `ir_patch_gen.py` to merge node attributes and parse additional fields. - Implemented `xls_ir_to_graph.cc` and `xls_ir_to_graph.h` for converting XLS IR to graph representation. - Added tests in `xls_ir_to_graph_test.cc` to validate graph conversion and debug node inclusion. - Cleaned up `patch_ir.cc` to handle assert and trace nodes directly in the IR graph.
- Introduced NodeCostAttributes and EdgeCostAttributes structs to encapsulate node and edge cost attributes, replacing string-based representations. - Updated XLSNode and XLSEdge to use the new attribute structs. - Modified the IR patch generation code to populate and utilize the new cost attribute structures. - Enhanced the xls_ir_to_graph conversion to construct NodeCostAttributes directly from IR nodes. - Updated tests to validate the new structure and ensure correct attribute handling.
…scheduled_main.cc
The Python diff/patch tools (ir_diff*, ir_patch_gen.py, xls_ir_to_*, etc.) and the Cytoscape exporter are superseded by the C++ MCS+GED pipeline; delete them and prune the BUILD file. Add README.md describing the diff -> apply -> verify chain, MCS/GED flags, logging, and references. Also rename linear_sum_assignment -> LinearSumAssignment to follow the Google C++ style for free functions.
Replace the ad-hoc std::chrono elapsed-time measurements in the GED and MCS
pipeline with xls::Stopwatch (xls/common/stopwatch.h), which wraps a steady
clock behind a single shared utility. Timing behaviour is unchanged.
- Convert always-on VLOG(0) statements to LOG(INFO), matching the rest of
XLS; VLOG levels 1-3 remain as verbosity tiers.
- Drop "using namespace xls" from patch_ir_main.cc in favour of explicit
xls:: qualification.
- Restore the lap_solver.h include guard (missing #endif).
- Record planned residual pruning, channel-aware patching, RTL build
automation, and benchmark expansion as TODO(xls-eco) markers at the
relevant sites.
The parser now uniquifies the state read's node name away from its state element (st -> st__1), so the hard-coded name no longer matches. Find the StateRead node in the parsed proc instead.
|
@proppy fixed |
This PR adds a C++ implementation of Graph Edit Distance (GED) for the ECO system, replacing the existing Python/NetworkX approach.
What's included
ged_main) and IR patch generationRegarding the new C++ GED
The Python version works but is slow on large graphs. This C++ implementation is faster, integrates directly with XLS IR APIs, and supports advanced features like pinned nodes and MCS optimization.
Migration
Python tools (ir_diff_main.py, ir2gxl.py) are marked for deprecation but still functional. They'll be removed in a future PR once the C++ version is proven stable.
Testing
Builds cleanly and includes tests with real DSLX examples (crc32, apfloat_fmac, riscv_simple, etc).
@grebe Please consider reviewing this PR. I tried to organize it into separate commits for convenience.