[SPARK-50127] Implement Avro encoding for MapState and PrefixKeyScanStateEncoder #22

ericm-db · 2024-11-04T03:31:27Z

What changes were proposed in this pull request?

Follow-up to apache#48401. This PR enables Avro encoding for MapState and the PrefixKeyScanStateEncoder

Why are the changes needed?

UnsafeRow is an inherently unstable format that makes no guarantees of being backwards-compatible. Therefore, if the format changes between Spark releases, this could cause StateStore corruptions. Avro is more stable, and inherently enables schema evolution.

Does this PR introduce any user-facing change?

No

How was this patch tested?

Amended and added to unit tests

Was this patch authored or co-authored using generative AI tooling?

No

* valuestatettl * mapstate, valuestate ttl works * timers * renaming to suffix key * cleaning up

…pressions in `buildAggExprList` ### What changes were proposed in this pull request? Trim aliases before matching Sort/Having/Filter expressions with semantically equal expression from the Aggregate below in `buildAggExprList` ### Why are the changes needed? For a query like: ``` SELECT course, year, GROUPING(course) FROM courseSales GROUP BY CUBE(course, year) ORDER BY GROUPING(course) ``` Plan after `ResolveReferences` and before `ResolveAggregateFunctions` looks like: ``` !Sort [cast((shiftright(tempresolvedcolumn(spark_grouping_id#18L, spark_grouping_id, false), 1) & 1) as tinyint) AS grouping(course)#22 ASC NULLS FIRST], true +- Aggregate [course#19, year#20, spark_grouping_id#18L], [course#19, year#20, cast((shiftright(spark_grouping_id#18L, 1) & 1) as tinyint) AS grouping(course)#21 AS grouping(course)#15] .... ``` Because aggregate list has `Alias(Alias(cast((shiftright(spark_grouping_id#18L, 1) & 1) as tinyint))` expression from `SortOrder` won't get matched as semantically equal and it will result in adding an unnecessary `Project`. By stripping inner aliases from aggregate list (that are going to get removed anyways in `CleanupAliases`) we can match `SortOrder` expression and resolve it as `grouping(course)#15` ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? Existing tests ### Was this patch authored or co-authored using generative AI tooling? No Closes apache#51339 from mihailotim-db/mihailotim-db/fix_inner_aliases_semi_structured. Authored-by: Mihailo Timotic <[email protected]> Signed-off-by: Wenchen Fan <[email protected]>

github-actions bot added SQL STRUCTURED STREAMING labels Nov 4, 2024

ericm-db force-pushed the avro-ps branch from 5384d86 to 0d53d0a Compare November 4, 2024 18:46

ericm-db changed the title ~~[WIP] Supporting Map State with Avro encoding~~ [SPARK-50127] Implement Avro encoding for MapState and PrefixKeyScanStateEncoder Nov 4, 2024

[WIP] Supporting Map State with Avro encoding

7c7f9db

ericm-db force-pushed the avro-ps branch from 0d53d0a to 7c7f9db Compare November 4, 2024 22:38

ericm-db and others added 3 commits November 4, 2024 14:45

adding comments

8f7a8c2

merge into origin/avro

9855130

[WIP] Avro Range Scan (#23)

0ed0ca0

* valuestatettl * mapstate, valuestate ttl works * timers * renaming to suffix key * cleaning up

ericm-db merged commit 9b8dd5d into avro Nov 7, 2024
3 of 4 checks passed

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Uh oh!

[SPARK-50127] Implement Avro encoding for MapState and PrefixKeyScanStateEncoder #22

[SPARK-50127] Implement Avro encoding for MapState and PrefixKeyScanStateEncoder #22

Uh oh!

ericm-db commented Nov 4, 2024 •

edited

Loading

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

1 participant

Uh oh!

[SPARK-50127] Implement Avro encoding for MapState and PrefixKeyScanStateEncoder #22

[SPARK-50127] Implement Avro encoding for MapState and PrefixKeyScanStateEncoder #22

Uh oh!

Conversation

ericm-db commented Nov 4, 2024 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

What changes were proposed in this pull request?

Why are the changes needed?

Does this PR introduce any user-facing change?

How was this patch tested?

Was this patch authored or co-authored using generative AI tooling?

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

1 participant

ericm-db commented Nov 4, 2024 •

edited

Loading