# 散列表 本节围绕以下内容展开: * 散列表 * 散列函数设计 * 冲突处理 * hashmap数据结构 * Java 中HashMap实现 散列表使用某种算法操作(散列函数)将键转化为数组的索引来访问数组中的数据,这样可以通过Key-value的方式来访问数据,达到常数级别的存取效率。现在的nosql数据库都是采用key-value的方式来访问存储数据。 散列表是算法在时间和空间上做出权衡的经典例子。通过一个散列函数,将键值key映射到记录的访问地址,达到快速查找的目的。如果没有内存限制,我们可以直接将键作为数组的索引,所有的操作操作只需要一次访问内存就可以完成。但这种情况不太现实。 ## Hashmap应用 1. cocos2d 游戏引擎 CCScheduler 2. linux 内核bcache。 缓存加速技术,使用SSD固态硬盘作为高速缓存,提高慢速存储设备HDD机械硬盘的性能 3. hash表在海量数据处理中有广泛应用。如海量日志中,提取出某日访问百度次数最多的IP 4. Java 中HashMap实现。编程语言中HashMap是如何实现的呢? 说说 Java , Golang 5. redis hash结构, set通常也是基于Hash结构实现 ## 散列函数 散列函数就是将键转化为数组索引的过程。且这个函数应该易于计算且能够均与分布所有的键。 散列函数最常用的方法是`除留余数法`。这时候被除数应该选用`素数`,这样才能保证键值的均匀散步。 散列函数和键的类型有关,每种数据类型都需要相应的散列函数;比如键的类型是整数,那我们可以直接使用`除留余数法`;这里特别说明下,大多数情况下,键的类型都是字符串,这个时候我们任然可以使用`除留余数法`,将字符串当做一个特别大的整数。 ``` int hash = 0; for (int i=0;isize = size; hashMap->usage = 0; hashMap->heads = calloc(size,sizeof(Entry *)); return hashMap; } HashMap *put(HashMap *hashMap,Key key,Value value){ if (key == NULL){ return hashMap; } Hash hash = hashCode(key); int index = hash & (size-1);/* */ if (hashMap->heads[index] == NULL){ _putInHead(hashMap,index,key,value); }else{ _putInList(hashMap,index,key,value); } } Value get(HashMap hashMap*,Key key){ if (key == NULL){ return hashMap; } Hash hash = hashCode(key); int index = hash & (size-1);/* */ Entry *entry = hashMap->heads[index]; while(entry != NULL){ if (entry->hash == hash){ return entry->value; } entry = entry->next; } return NULL; } HashMap *_putInHead(HashMap *hashMap,int index,Key key,Value value){ Entry *newHead = malloc(sizeof(Entry)); newHead->hash = hash; newHead->key = key; newHead->value = value; hashMap->heads[index] = newHead; (hashMap->usage)++; return hashMap; } HashMap *_putInList(HashMap *hashMap,int index,Key key,Value value){ Entry *lastEntry = hashMap->heads[index]; while(lastEntry != NULL){ if (lastEntry->hash == hash){ return hashMap; }else{ lastEntry = lastEntry->next; } } lastEntry = malloc(sizeof(Entry)); lastEntry->hash = hash; lastEntry->key = key; lastEntry->value = value; lastEntry->next = hashMap->heads[index]; hashMap->heads[index] = lastEntry; (hashMap->usage)++; return hashMap; } ``` ### 扩容 当hash表保存的键值对数量太多或太少,对hash 表进行扩容和缩容。合理控制内存的使用。 redis hash中, rehash发生在扩容或缩容阶段,扩容是发生在元素的个数等于哈希表数组的长度时,进行2倍的扩容;缩容发生在当元素个数为数组长度的10%时,进行缩容 ## 参考 《Algorithms》