2017-07-18 65 views
2

我在HBase中有一个表,其中汉字存储在某个列中,比如'FLT:CREW_DEPT'。现在我需要过滤出'FLT:CREW_DEPT'等于某个值的所有行。当HBase的壳这样做,它工作正常,如下图所示当使用happybase扫描一个带有中文字符的HBase表时,过滤器不起作用

scan 'PAX_EXP_FACT', {COLUMNS => ['FLT:CREW_DEPT'], FILTER => "SingleColumnValueFilter ('FLT', 'CREW_DEPT', =, 'binary:\xe4\xb8\x80\xe9\x83\xa8', true, true)", LIMIT => 5} 
ROW             COLUMN+CELL 
CA101-20160808-PEK-001192753702      column=FLT:CREW_DEPT, timestamp=1500346136328, value=\xE4\xB8\x80\xE9\x83\xA8 
CA101-20161103-PEK-001181988752      column=FLT:CREW_DEPT, timestamp=1500346230204, value=\xE4\xB8\x80\xE9\x83\xA8 
CA101-20161105-PEK-000728690130      column=FLT:CREW_DEPT, timestamp=1500346244963, value=\xE4\xB8\x80\xE9\x83\xA8 
CA101-20161201-PEK-006731936575      column=FLT:CREW_DEPT, timestamp=1500346233640, value=\xE4\xB8\x80\xE9\x83\xA8 
CA101-20161212-PEK-001512808262      column=FLT:CREW_DEPT, timestamp=1500346223572, value=\xE4\xB8\x80\xE9\x83\xA8 
5 row(s) in 0.0060 seconds 

然而,使用happybase在Python做类似的事情的时候,什么也没有返回:返回

import happybase 
import datetime 
import pytz 

connection = happybase.Connection('192.168.199.200', port=9090) 
table = connection.table('PAX_EXP_FACT') 

filter_str = "" 
filter_str += "SingleColumnValueFilter('FLT', 'CREW_DEPT', =, 'binary:\xe4\xba\x8c\xe9\x83\xa8')" 

results = table.scan(
    filter=filter_str 
    #  ,limit=100 
) 

count = 0 
for key, data in results: 
    count += 1 
    print(data[b'FLT:CREW_DEPT'].decode('utf-8')) 

print('No. of flight matches:', count) 

connection.close() 

0行...

任何人都可以帮忙吗?非常感激!!!

回答

1

答案竟然是非常简单的......我应该用

"SingleColumnValueFilter('FLT', 'CREW_DEPT', =, 'binary:中文')" 

,而不是首先将其转换为UTF-8编码的字节串......即使在HBase的壳,我可以做同样的事情(尽管它显示在问号)

scan 'PAX_EXP_FACT', {COLUMNS => ['FLT:CREW_DEPT'], FILTER => "SingleColumnValueFilter ('FLT', 'CREW_DEPT', =, 'binary:??', true, true)", LIMIT => 5} 
ROW             COLUMN+CELL 
CA101-20160808-PEK-001192753702      column=FLT:CREW_DEPT, timestamp=1500353334419, value=\xE4\xB8\x80\xE9\x83\xA8 
CA101-20161103-PEK-001181988752      column=FLT:CREW_DEPT, timestamp=1500353426641, value=\xE4\xB8\x80\xE9\x83\xA8 
CA101-20161105-PEK-000728690130      column=FLT:CREW_DEPT, timestamp=1500353447707, value=\xE4\xB8\x80\xE9\x83\xA8 
CA101-20161201-PEK-006731936575      column=FLT:CREW_DEPT, timestamp=1500353432222, value=\xE4\xB8\x80\xE9\x83\xA8 
CA101-20161212-PEK-001512808262      column=FLT:CREW_DEPT, timestamp=1500353417107, value=\xE4\xB8\x80\xE9\x83\xA8 
5 row(s) in 0.0100 seconds 

使用子就不行,这是我一直在努力之前,我提出了这个愚蠢的问题...