Robert Howell
01/21/2025, 7:09 PMDESCRIBE EXTENDED is used to describe relational level statistics rather than column level statistics which suggests it is not the correct fit for the summarizing column statistics.
Option 1 is to my personal taste, and is also technically simpler. Desmond has already PR'd a command for returning column statistics, while I have been working on the SQL DESCRIBE. Some more details in the 🧵Robert Howell
01/21/2025, 7:09 PMprint(df.schema()) will ultimately call the Display trait of Schema which makes a pretty-printed table. Rather than having DESCRIBE just return this string, we want it to return a relation.
// <http://PySchema.rs|PySchema.rs> <- _PySchema.py <- Schema.py <- df.schema()
impl PySchema {
// ...
pub fn __repr__(&self) -> PyResult<String> {
Ok(format!("{}", self.schema))
}
}
// Display for Schema
#[display("{}\n", make_schema_vertical_table(
fields.iter().map(|(name, field)| (name.clone(), field.dtype.to_string()))
))]
impl Schema { }Robert Howell
01/21/2025, 7:17 PMDESCRIBE TABLE tbl;
-- output schema
-- +--------------+------------+
-- | column(text) | type(text) |
-- +--------------+------------+
DESCRIBE (SELECT * FROM 'data.csv')
usage
DESCRIBE [TABLE] <table>; -- df.describe()
DESCRIBE <query>; -- df.describe()
SUMMARIZE <table>; -- df.summarize()Desmond Cheong
01/21/2025, 7:18 PMDESCRIBE Statement produces a relation representing the schema of the relation being described.Yeah I'm a bigger fan of this than shoving stats into the describe. Seems pretty standard to put name + type (+ nullable + default + collation + key info (?)) Hadn't done a competitor analysis when I wrote the PR to display stats (I was a little influenced by DBR's DESCRIBE EXTENDED, and spark syntax in general). Fine with SUMMARIZE
Robert Howell
01/21/2025, 7:19 PMSeems pretty standard to put name + type (+ nullable + default + collation + key info (?))yes, coming from SQL spec. I notices DBR's (and others) describe extended include additional rows rather than additional columns like SUMMARIZE does. • Supports DESCRIBE with an “EXTENDED” to return additional metadata about the relation and is a slightly different statement than the SQL DESCRIBE. • The documentation suggests this includes “additional information” like column stats, but the example outputs look different. The DESCRIBE EXTENDED “Returns additional metadata such as parent schema, owner, access time etc." • The Databricks DESCRIBE EXTENDED actually includes more ROWS rather than more COLUMNS. We should prefer to include more columns because (1) it demarcates the schema columns (which become rows in describe) from metadata information and (2) the SQL standard makes room for additional output columns of DESCRIBE so we are actually better aligning with the standard here. • Databricks has ANALYZE https://docs.databricks.com/en/sql/language-manual/sql-ref-syntax-aux-analyze-table.html which is a bit more similar and aligns with PostgreSQL to some degree.
Robert Howell
01/21/2025, 7:21 PMdf.summarize() # 1
# OR
df.describe(summarize=True) # 2
I'm inclined to #1Desmond Cheong
01/21/2025, 7:22 PMANALYZE is a little different because it also enriches the catalog with statistics for the table (or refreshes stats if they are out of date). I'd say that's a different class of commandRobert Howell
01/21/2025, 7:22 PMDesmond Cheong
01/21/2025, 7:24 PMI'm inclined to #1+1
Robert Howell
01/21/2025, 7:29 PM