From 3a88447c9d3a2386dc64d11b87f8834223a69f93 Mon Sep 17 00:00:00 2001 From: ranwenjie Date: Sat, 6 Nov 2021 23:46:45 +0800 Subject: [PATCH] update readme --- 3 Hash Table/README.md | 1 + 5 Graph/DFS 和 BFS.md | 5 +++++ 6 Sort/外排序.md | 5 +++++ .../{bloomfilter => }/Bitmap.md | 12 ++++++------ ...loom filter(布隆过滤器).md => Bloomfilter.md} | 1 - .../{mapreduce => }/Hash映射,分而治之.md | 0 .../Inverted Index/倒排索引(Inverted Index).md | 3 +-- .../Inverted Index/数据库索引.md | 3 +-- 91 Algorithms In Big Data/README.md | 16 ++++++++-------- .../{simhash => }/simhash算法.md | 0 91 Algorithms In Big Data/双层桶划分.md | 4 +--- 92 Algorithms In DB/README.md | 3 +-- 12 files changed, 29 insertions(+), 24 deletions(-) create mode 100644 6 Sort/外排序.md rename 91 Algorithms In Big Data/{bloomfilter => }/Bitmap.md (78%) rename 91 Algorithms In Big Data/{bloomfilter/Bloom filter(布隆过滤器).md => Bloomfilter.md} (99%) rename 91 Algorithms In Big Data/{mapreduce => }/Hash映射,分而治之.md (100%) rename 91 Algorithms In Big Data/{simhash => }/simhash算法.md (100%) diff --git a/3 Hash Table/README.md b/3 Hash Table/README.md index 409a780..0bf1d4f 100644 --- a/3 Hash Table/README.md +++ b/3 Hash Table/README.md @@ -63,6 +63,7 @@ Hash hashCode(char *key){ 当然,还有其他的散列函数,如`平方取中法`, `随机数法`等。 + ## 碰撞解决 diff --git a/5 Graph/DFS 和 BFS.md b/5 Graph/DFS 和 BFS.md index ecd990f..ed7d73c 100644 --- a/5 Graph/DFS 和 BFS.md +++ b/5 Graph/DFS 和 BFS.md @@ -1,5 +1,10 @@ # DFS 和 BFS 搜索算法 +DFS: 深度优先搜索,以深度为准则,先一条路走到底,直到达到目标; 没有达到目标又无路可走了,那么则退回到上一步的状态,走其他路。这便是回溯上来。 + +BFS:广度优先搜素,在面临一个路口时,把所有的岔路口都记下来,然后选择其中一个进入,然后将它的分路情况记录下来,然后再返回来进入另外一个岔路,并重复这样的操作。 + + > DFS用递归的形式,用到了栈结构,先进后出; BFS选取状态用队列的形式,先进先出。 diff --git a/6 Sort/外排序.md b/6 Sort/外排序.md new file mode 100644 index 0000000..6c426c6 --- /dev/null +++ b/6 Sort/外排序.md @@ -0,0 +1,5 @@ +# 外排序 + + + + diff --git a/91 Algorithms In Big Data/bloomfilter/Bitmap.md b/91 Algorithms In Big Data/Bitmap.md similarity index 78% rename from 91 Algorithms In Big Data/bloomfilter/Bitmap.md rename to 91 Algorithms In Big Data/Bitmap.md index d2ce2d3..94ac1f4 100644 --- a/91 Algorithms In Big Data/bloomfilter/Bitmap.md +++ b/91 Algorithms In Big Data/Bitmap.md @@ -1,16 +1,16 @@ - - -## Bitmap +# Bitmap 也就是用1个(或几个)bit位来标记某个元素对应的value(如果是1bitmap,就只能是元素是否存在;如果是x-bitmap,还可以是元素出现的次数等信息)。使用bit位来存储信息,在需要的存储空间方面可以大大节省。应用场景有: -1. 排序(如果是1-bitmap,就只能对无重复的数排序) -2. 判断某个元素是否存在 +1. 判断某个元素是否存在 +2. 排序(如果是1-bitmap,就只能对无重复的数排序) + 比如,某文件中有若干8位数字的电话号码,要求统计一共有多少个不同的电话号码? -分析:8位最多99 999 999, 如果1Byte表示1个号码是否存在,需要95MB空间,但是如果1bit表示1个号码是否存在,则只需要 95/8=12MB 的空间。这时,数字k(0~99 999 999)与bit位的对应关系是: +分析:8位最多99 999 999, 如果1Byte表示1个号码是否存在,需要95MB空间,但是如果1bit表示1个号码是否存在,则只需要 95/8=12MB 的空间。这时,数字 `k(0~99 999 999)`与bit位的对应关系是: + ``` #define SIZE 15*1024*1024 diff --git a/91 Algorithms In Big Data/bloomfilter/Bloom filter(布隆过滤器).md b/91 Algorithms In Big Data/Bloomfilter.md similarity index 99% rename from 91 Algorithms In Big Data/bloomfilter/Bloom filter(布隆过滤器).md rename to 91 Algorithms In Big Data/Bloomfilter.md index fcd6520..c0075aa 100644 --- a/91 Algorithms In Big Data/bloomfilter/Bloom filter(布隆过滤器).md +++ b/91 Algorithms In Big Data/Bloomfilter.md @@ -1,4 +1,3 @@ - # Bloom filter(布隆过滤器) Bloom Filter是由Bloom在1970年提出的一种多哈希函数映射的快速查找算法。通常应用在海量数据处理中,一些需要快速判断某个元素是否属于集合,但是并不严格要求100%正确的场合(容忍错误的场景)。 diff --git a/91 Algorithms In Big Data/mapreduce/Hash映射,分而治之.md b/91 Algorithms In Big Data/Hash映射,分而治之.md similarity index 100% rename from 91 Algorithms In Big Data/mapreduce/Hash映射,分而治之.md rename to 91 Algorithms In Big Data/Hash映射,分而治之.md diff --git a/91 Algorithms In Big Data/Inverted Index/倒排索引(Inverted Index).md b/91 Algorithms In Big Data/Inverted Index/倒排索引(Inverted Index).md index aed743c..b203143 100644 --- a/91 Algorithms In Big Data/Inverted Index/倒排索引(Inverted Index).md +++ b/91 Algorithms In Big Data/Inverted Index/倒排索引(Inverted Index).md @@ -1,5 +1,4 @@ - -## 倒排索引(Inverted Index) +# 倒排索引(Inverted Index) 常规的索引是文档到关键词的映射,就是每个文档指向一个它所包含的索引项的序列,也就是文档文档指向了它包含的索引项序列,也就是文档指向它包含的哪些单词。 diff --git a/91 Algorithms In Big Data/Inverted Index/数据库索引.md b/91 Algorithms In Big Data/Inverted Index/数据库索引.md index 48048f7..96b8cf0 100644 --- a/91 Algorithms In Big Data/Inverted Index/数据库索引.md +++ b/91 Algorithms In Big Data/Inverted Index/数据库索引.md @@ -1,5 +1,4 @@ - -## 数据库索引 +# 数据库索引 索引使用的数据结构多是B树或B+树。B树和B+树广泛应用于文件存储系统和数据库系统中,mysql使用的是B+树,oracle使用的是B树,Mysql也支持多种索引类型,如b-tree 索引,哈希索引,全文索引等。 diff --git a/91 Algorithms In Big Data/README.md b/91 Algorithms In Big Data/README.md index 9a34953..b45e6b5 100644 --- a/91 Algorithms In Big Data/README.md +++ b/91 Algorithms In Big Data/README.md @@ -6,15 +6,15 @@ 针对空间,就一个办法,大而化小,分而治之。常采用hash映射 -* Hash映射/分而治之 -* Bitmap -* Bloom filter(布隆过滤器) -* 双层桶划分 -* Trie树 -* 数据库索引 +* [Hash映射,分而治之](Hash映射,分而治之.md) +* [Bitmap](Bitmap.md) +* [Bloom filter(布隆过滤器)](Bloomfilter.md) +* [双层桶划分](双层桶划分.md) +* [Trie树](../4%20Tree/4-字典树Trie/README.md) +* [数据库索引](Inverted%20Index/数据库索引.md) * 倒排索引(Inverted Index) -* 外排序 -* simhash算法 +* [外排序](../6%20Sort/外排序.md) +* [simhash算法](simhash算法.md) * 分布处理之Mapreduce diff --git a/91 Algorithms In Big Data/simhash/simhash算法.md b/91 Algorithms In Big Data/simhash算法.md similarity index 100% rename from 91 Algorithms In Big Data/simhash/simhash算法.md rename to 91 Algorithms In Big Data/simhash算法.md diff --git a/91 Algorithms In Big Data/双层桶划分.md b/91 Algorithms In Big Data/双层桶划分.md index 01d77be..c4adf7d 100644 --- a/91 Algorithms In Big Data/双层桶划分.md +++ b/91 Algorithms In Big Data/双层桶划分.md @@ -1,6 +1,4 @@ - - -## 双层桶划分 +# 双层桶划分 双层桶不是一种数据结构,只是一种算法思维。分而治之思想。 diff --git a/92 Algorithms In DB/README.md b/92 Algorithms In DB/README.md index 4095714..f818b84 100644 --- a/92 Algorithms In DB/README.md +++ b/92 Algorithms In DB/README.md @@ -1,5 +1,4 @@ - -## 数据库系统中的算法 +# 数据库系统中的算法 最近开始读《数据库系统实现》这本书,所以就想到把数据库里面用到的数据结构和算法做一个梳理。就有了这些文字。