我在google云上有一个Mysql数据库,我想创建自动模式将数据插入Bigquery,我需要自动创建以下行:
schema= [bigquery.SchemaField('EmployeeID', 'STRING', mode='NULLABLE')
bigquery.SchemaField('LastName', 'STRING', mode='NULLABLE')
bigquery.SchemaField('FirstName', 'STRING', mode='NULLABLE')
bigquery.SchemaField('Title', 'STRING', mode='NULLABLE')
bigquery.SchemaField('TitleOfCourtesy', 'STRING', mode='NULLABLE')
bigquery.SchemaField('BirthDate', 'STRING', mode='NULLABLE')
bigquery.SchemaField('HireDate', 'STRING', mode='NULLABLE')
bigquery.SchemaField('Address', 'STRING', mode='NULLABLE')
bigquery.SchemaField('City', 'STRING', mode='NULLABLE')
bigquery.SchemaField('Region', 'STRING', mode='NULLABLE')
bigquery.SchemaField('PostalCode', 'STRING', mode='NULLABLE')
bigquery.SchemaField('Country', 'STRING', mode='NULLABLE')
bigquery.SchemaField('HomePhone', 'STRING', mode='NULLABLE')
bigquery.SchemaField('Extension', 'STRING', mode='NULLABLE')
bigquery.SchemaField('Photo', 'STRING', mode='NULLABLE')
bigquery.SchemaField('Notes', 'STRING', mode='NULLABLE')
bigquery.SchemaField('ReportsTo', 'STRING', mode='NULLABLE')
bigquery.SchemaField('PhotoPath', 'STRING', mode='NULLABLE')]
所以为了实现这一点,我试过:首先,我得到一个带有函数的列的名称,这是我的输出:
print(table_schema_name_column)
['EmployeeID', 'LastName', 'FirstName', 'Title', 'TitleOfCourtesy', 'BirthDate', 'HireDate', 'Address', 'City', 'Region', 'PostalCode', 'Country', 'HomePhone', 'Extension', 'Photo', 'Notes', 'ReportsTo', 'PhotoPath']
然后我试过:
schema2=[]
for element in table_schema_name_column:
base2="bigquery.SchemaField("+'\''+element+"\', \'STRING\', mode=\'NULLABLE\')"
tmp=base2
#print(base2)
schema2.append(base2)
print(schema2)
这是相应的输出:
["bigquery.SchemaField('EmployeeID', 'STRING', mode='NULLABLE')",
"bigquery.SchemaField('LastName', 'STRING', mode='NULLABLE')",
"bigquery.SchemaField('FirstName', 'STRING', mode='NULLABLE')", "bigquery.SchemaField('Title', 'STRING', mode='NULLABLE')",
"bigquery.SchemaField('TitleOfCourtesy', 'STRING', mode='NULLABLE')", "bigquery.SchemaField('BirthDate', 'STRING', mode='NULLABLE')",
"bigquery.SchemaField('HireDate', 'STRING', mode='NULLABLE')", "bigquery.SchemaField('Address', 'STRING', mode='NULLABLE')",
"bigquery.SchemaField('City', 'STRING', mode='NULLABLE')", "bigquery.SchemaField('Region', 'STRING', mode='NULLABLE')",
"bigquery.SchemaField('PostalCode', 'STRING', mode='NULLABLE')", "bigquery.SchemaField('Country', 'STRING', mode='NULLABLE')",
"bigquery.SchemaField('HomePhone', 'STRING', mode='NULLABLE')", "bigquery.SchemaField('Extension', 'STRING', mode='NULLABLE')",
"bigquery.SchemaField('Photo', 'STRING', mode='NULLABLE')", "bigquery.SchemaField('Notes', 'STRING', mode='NULLABLE')",
"bigquery.SchemaField('ReportsTo', 'STRING', mode='NULLABLE')", "bigquery.SchemaField('PhotoPath', 'STRING', mode='NULLABLE')"]
这个schema2的问题是,当我尝试使用它来创建表时,我收到以下错误:
table_ref = dataset_ref.table("my_table_aut")
table = bigquery.Table(table_ref, schema=schema2)
table = client.create_table(table) # API request
assert table.table_id == "my_table_aut"
错误输出:
ValueError Traceback (most recent call last)
<ipython-input-13-ce1fc2c637fe> in <module>
4 ]
5 table_ref = dataset_ref.table("my_table_aut")
----> 6 table = bigquery.Table(table_ref, schema=schema2)
7 table = client.create_table(table) # API request
8
~/.local/lib/python3.6/site-packages/google/cloud/bigquery/table.py in __init__(self, table_ref, schema)
371 # Let the @property do validation.
372 if schema is not None:
--> 373 self.schema = schema
374
375 @property
~/.local/lib/python3.6/site-packages/google/cloud/bigquery/table.py in schema(self, value)
420 self._properties["schema"] = None
421 elif not all(isinstance(field, SchemaField) for field in value):
--> 422 raise ValueError("Schema items must be fields")
423 else:
424 self._properties["schema"] = {"fields": _build_schema_resource(value)}
ValueError: Schema items must be fields
所以我想感谢支持克服这一任务
这应该工作:
schema2=[]
for element in table_schema_name_column:
schema2.append(bigquery.SchemaField(element, 'STRING', mode='NULLABLE'))
table_ref = dataset_ref.table("my_table_aut")
table = bigquery.Table(table_ref, schema=schema2)
table = client.create_table(table)