Adds aggregation/bhr_collection.py with a streaming get_data() generator. The SQL, FARM_FINGERPRINT-based sampling, and target table are unchanged from the python_mozetl original. Streaming keeps memory bounded. At sample-size 0.5 the result set is roughly 500K rows of 5-10 KB each, which would need 3-5 GB if materialised into a list. The google-cloud-bigquery import is lazy so the module imports cleanly without the package installed (unit tests mock the client). The production runtime will install the package via TaskCluster's Docker image in a later phase. Differential Revision: https://phabricator.services.mozilla.com/D303370
11 lines
149 B
TOML
11 lines
149 B
TOML
[DEFAULT]
|
|
subsuite = "bhr-aggregation"
|
|
|
|
["test_bhr_collection.py"]
|
|
|
|
["test_heuristics.py"]
|
|
|
|
["test_profile_processor.py"]
|
|
|
|
["test_symbolication.py"]
|